← Reddit

Claude is really bad at analyzing writing, but it gives such confident analyses that it's easy to miss just how bad it is

Reddit · RampantInanity · July 31, 2026
A master's student and teacher testing Claude for writing analysis found that despite its confident tone, the system frequently missed overarching meaning and context while fixating on minor points and misinterpreting complete sentences. In one instance, Claude gave harsh feedback on a sentence based solely on its second half, despite the full sentence conveying a meaningfully different meaning. The experience revealed that LLMs remain unreliable analytical tools for critical tasks like student grading, as their authoritative presentation can mislead users unfamiliar with the content.

Detailed Analysis

A Reddit post from a master's student and teacher has drawn attention to a persistent weakness in Claude's capabilities: its unreliability as an analytical tool for evaluating writing, despite presenting its judgments with unwarranted confidence. The author, who uses Claude regularly for other tasks, tested the model on essays and found that it consistently misread the relationship between thesis statements and supporting paragraphs, fixated on minor details while missing broader argumentative structure, and in one striking example, claimed a sentence "detonated" the thesis when in fact only half the sentence supported that reading — the full sentence conveyed something quite different. The documents involved were modest in length, around 50 pages, well within Claude's advertised context window, yet the model still failed to track the coherence of the argument across the text.

This gap between apparent comprehension and actual comprehension is not a minor technical footnote — it cuts to one of the central critiques of large language models generally: that they are exceptionally good at producing fluent, authoritative-sounding text without possessing the kind of holistic understanding that task requires. A large context window means a model can technically process more tokens, but as this example shows, ingesting text is not the same as synthesizing it. Claude can retrieve and reference passages, but linking a claim on page one to a complication on page thirty in a way that reflects true argumentative coherence is a different and much harder problem, one current transformer-based architectures still struggle with.

The stakes here are practical and immediate. Teachers increasingly use Claude and similar tools to assist with grading, feedback generation, and instructional support, often trusting that a model's polished, confident tone reflects sound underlying analysis. The Reddit author's account is a cautionary tale specifically because the errors were subtle rather than glaring nonsense — a superficially plausible critique built on a misreading that would be invisible to anyone not already deeply familiar with the text. That is arguably more dangerous than an obvious hallucination: a student handed inaccurate, harshly worded feedback about a nonexistent flaw might waste time "fixing" something that was never broken, or lose trust in feedback that happened to be accurate elsewhere.

More broadly, this incident reinforces a distinction increasingly drawn by experienced AI users: LLMs tend to excel at generative or transformative tasks within a domain the user already understands — drafting, summarizing, rephrasing, brainstorming — but remain far less trustworthy as independent evaluators or analysts of material outside the user's own expertise, precisely because their confident tone doesn't correlate with accuracy. As adoption of tools like Claude accelerates in education, law, medicine, and other fields where nuanced judgment matters, this tension between fluency and genuine reasoning is likely to remain one of the most consequential open problems in applied AI, and one that model providers like Anthropic will need to address as they push Claude toward more autonomous, judgment-intensive use cases.

Read original article →