Detailed Analysis
This Reddit post highlights a recurring pain point for Claude.ai subscribers: the difficulty of predicting and controlling usage costs when interacting with Claude through custom personas or extended-thinking configurations. The user, a Pro plan subscriber who primarily runs text-based analysis prompts, describes hitting usage limits while using "Fable 5"—apparently a custom prompt, project, or character configuration built on Claude that leverages Claude's extended thinking feature. Their core complaint is straightforward but important: extended thinking generates a "thinking block" that is often longer than the visible answer itself, yet this reasoning text consumes tokens and counts against usage limits or costs, even though users don't directly see or control its length. The post reflects genuine confusion about how to estimate per-prompt costs when a significant, variable-length portion of the output (the model's internal reasoning) is effectively invisible until after the fact.
This matters because Anthropic's extended thinking mode, while valuable for improving answer quality on complex reasoning tasks, introduces a transparency and predictability problem for cost-conscious users. Unlike traditional token-based API pricing where developers can estimate costs by counting input and output tokens in advance, extended thinking adds a dynamic, model-determined amount of "thinking" tokens that aren't easily forecasted. For Pro plan users on Claude.ai, this is compounded by the fact that Anthropic uses usage caps (rather than direct per-token billing) for consumer-tier access, meaning users experience the cost indirectly through hitting session or daily limits rather than seeing a running dollar total. When a single extended-thinking response can be dominated by a "thinking" block longer than the actual answer, users have no reliable way to budget their remaining quota across a session, especially when using third-party or community-built characters/prompts (like "Fable") that may enable extended thinking by default without the user having full control over that toggle.
The timing reference—"after today"—suggests this post coincides with a change to Anthropic's usage policies, pricing, or plan limits taking effect, which is a common source of anxiety in Anthropic's user community. Anthropic has periodically adjusted Pro and Max plan usage limits, rate-limiting windows, and the way extended thinking factors into quota consumption, and such changes often arrive with limited advance communication, leaving power users scrambling to understand new constraints. This is a recurring theme across r/Anthropic and r/ClaudeAI: users build workflows or rely on specific custom GPT-like configurations, then find those workflows disrupted or made unpredictable by backend changes to token accounting, thinking budgets, or rate limits.
More broadly, this incident is emblematic of a tension running through the entire generative AI industry: as models gain more sophisticated reasoning capabilities (extended/chain-of-thought thinking), the computational and financial cost of those capabilities becomes harder for end users to reason about, ironically mirroring the opacity of the reasoning process itself. Competitors like OpenAI (with o1/o3 reasoning models) face the same criticism—reasoning tokens are billed but not always fully visible, creating a cost/UX mismatch. For consumer-facing products like Claude.ai, where users pay flat subscription fees rather than per-token API rates, this opacity is even more acute, since there's no direct way to translate "thinking tokens used" into "dollars spent" the way developers can on the API. Anthropic will likely face continued pressure to provide better usage dashboards, thinking-budget controls, or clearer documentation so that everyday subscribers—not just API developers—can understand and manage the true cost of increasingly capable but computationally expensive reasoning features.
Read original article →