← Reddit

We hear a lot of what Claude can do. What is Claude not able to do... yet?

Reddit · Eurofan4640 · August 1, 2026

Detailed Analysis

The Reddit thread in question surfaces a recurring theme in the Claude user community: despite the rapid pace of capability announcements from Anthropic, everyday practitioners continue to encounter meaningful limitations when applying Claude to large, real-world projects. The original poster's question—framed around "full-sized reports and projects"—points to a gap between benchmark performance and the messier demands of professional workflows, where documents span dozens of pages, require sustained internal consistency, and depend on context that accumulates over hours or days of work rather than a single prompt-response exchange.

This gap matters because it highlights the difference between narrow task competence and genuine workflow integration. Claude models have demonstrated strong performance on coding benchmarks, reasoning tasks, and shorter-form writing, but users report friction points such as maintaining coherent structure across long documents, remembering earlier decisions or stylistic choices without explicit re-prompting, handling complex multi-file or multi-source synthesis, and avoiding drift or repetition in extended outputs. Context window size, while dramatically expanded in recent Claude versions (including 200K-token and larger windows), does not automatically translate into perfect recall or consistent judgment across that entire span—models can still lose track of earlier constraints, contradict prior statements, or fail to apply a consistent voice throughout a lengthy deliverable. These are the kinds of practical shortcomings that matter enormously to knowledge workers producing business reports, research documents, legal briefs, or technical specifications, even if they don't show up prominently in standard evaluation suites.

The broader significance of this kind of community discussion is that it functions as an informal, crowdsourced audit of AI capability claims. Marketing materials and official benchmarks tend to emphasize peak performance under favorable conditions, while forums like r/ClaudeAI aggregate the accumulated experience of thousands of daily users pushing models into edge cases and sustained real-world use. This feedback loop has become an important, if unofficial, signal for both users deciding how to allocate tasks between human and AI effort, and indirectly for developers like Anthropic who monitor such discussions to identify where models fall short of user expectations for "agentic" or long-horizon work.

This tension also reflects a larger trend across the AI industry: the shift in ambition from single-turn question answering toward autonomous or semi-autonomous execution of extended, multi-step projects—drafting entire reports, managing research pipelines, or acting as persistent collaborators. Anthropic's own product roadmap, including features like Projects, extended thinking modes, and computer-use capabilities, signals an explicit push toward exactly this kind of long-horizon reliability. Threads like this one serve as a reality check on that ambition, illustrating that while models like Claude have narrowed the gap considerably, sustained coherence, memory, and judgment over large-scale creative and analytical work remain an active frontier rather than a solved problem. The candid, specifics-seeking nature of the original post—asking for concrete experiences rather than general praise or criticism—also reflects a maturing user base that has moved past novelty and is now rigorously mapping the practical boundaries of what these tools can reliably do today, versus what remains aspirational.

Read original article →