← Reddit

Just replace Opus 5 with Fable 5 to be in the business. Which one is dumber even on Ultracode? Opus 4.8 Vs 5

Reddit · Regular_Attitude_700 · August 12, 2026
A customer reported that Claude Opus 5 on Anthropic's Ultracode platform performed poorly for their use case and failed to properly read project files. After switching to Fable 5 as an alternative, the model exhausted their monthly token quota within two hours of use, leaving them without service for five days until the quota reset. The customer expressed dissatisfaction with both models and issued an ultimatum that Anthropic replace Opus 5 or risk losing their business to competing AI services.

Detailed Analysis

A Reddit post to r/Anthropic captures a wave of frustration from a paying Claude Max subscriber ($200/month tier) who reports severe dissatisfaction with Opus 5's coding performance, particularly within what the poster calls "Ultracode" workflows. The user claims Opus 5 fails to properly read or follow project configuration files like claude.md, requires constant micromanagement, and performs worse than expected despite strong benchmark scores. The post's title references "Opus 4.8 vs 5," suggesting the user perceives a regression or inconsistency between recent model versions rather than the clear improvement Anthropic's release notes would imply. Compounding the frustration, the user describes an alternative or supplementary product referred to as "Fable 5" as consuming usage quotas at an unsustainable rate — draining a full week's allowance in roughly two of five planned work sessions, leaving them locked out for five days.

This complaint is emblematic of a recurring tension in the AI coding-assistant space: the gap between aggregate benchmark performance and real-world, task-specific reliability. Benchmarks like SWE-bench or various "agentic coding" leaderboards measure success on curated problem sets, but production coding workflows involve messy, project-specific context (custom instructions, large codebases, existing conventions) that models can fail to internalize consistently. When a model ignores a repository's guidance file or requires heavy hand-holding, the practical value proposition of "autonomous coding agent" erodes quickly, especially for professional users who are billing their own time against subscription costs. The user's explicit rejection of "harness, tool calling, and best practices" explanations reflects a broader frustration among power users who feel that Anthropic and similar vendors deflect legitimate product complaints onto user error or implementation details, rather than acknowledging model-level shortcomings.

The usage-quota complaint about aggressive token consumption on a rate-limited plan touches a separate but related pain point: as Anthropic (and competitors like OpenAI) push increasingly capable but computationally expensive models and features, they must balance capability against the sustainability of flat-rate subscription pricing. Power users running extended agentic sessions can exhaust weekly allowances rapidly, especially with models that consume more tokens through longer reasoning chains, more tool calls, or verbose outputs. This creates a friction point where the very users most likely to extract value from advanced coding agents — those running multi-hour, high-intensity sessions — are also the most likely to hit throttling limits, undermining retention among a company's most engaged and vocal customer segment.

More broadly, the post situates this individual grievance within larger anxieties about American AI leadership, referencing U.S. government efforts to regulate AI development pace and speculating that "Eastern AI companies" are gaining ground while Western firms face friction from both regulation and product quality issues. While hyperbolic, this reflects a genuine undercurrent in developer communities: as Chinese labs (DeepSeek, Alibaba's Qwen, Moonshot's Kimi, and others) ship increasingly competitive open-weight coding models at lower cost and with fewer usage restrictions, Western premium-subscription models like Claude face pressure to justify their pricing through consistent, superior real-world performance rather than benchmark leaderboard position alone. Posts like this — emotionally charged, anecdotal, but widely upvoted or discussed — function as informal signals to companies like Anthropic about churn risk, especially when frustration coalesces around specific failure modes (ignoring project context, inconsistent version-to-version quality, and quota exhaustion) that are fixable through engineering and product decisions rather than fundamental model limitations.

Read original article →