Detailed Analysis
A Reddit user on Anthropic's Max 5 subscription plan has raised concerns about a noticeable degradation in usable session capacity following the release of Opus 5, Anthropic's latest flagship model. The poster describes a workflow built around "Fable 5" orchestrating multiple sub-agents in extended, high-effort sessions—something that reportedly ran smoothly for hours at a time under the previous model generation (referred to as "4.8") without approaching plan limits. Since upgrading to Opus 5, the same user claims to be hitting both session-level and daily usage caps significantly faster, despite doing comparably less work. The post includes a screenshot purporting to show the user is nowhere near their monthly quota, which they present as evidence that the sudden constraint is not simply a matter of heavier personal usage but something structural or model-driven.
The complaint touches on a persistent tension in how AI companies price and meter access to increasingly capable models. Newer, more powerful models like Opus 5 often carry higher per-token computational costs due to larger parameter counts, more extensive reasoning chains, or increased context processing—even when the visible task output looks similar to what a smaller or older model produced. If Opus 5 is generating longer chains of thought, making more tool calls, or orchestrating sub-agents with greater verbosity per action, token consumption could rise substantially without a proportional increase in perceived work completed. This dynamic is especially acute for "agentic" workflows, where a single directive can cascade into dozens or hundreds of downstream model calls across sub-agents, compounding any per-call cost increase into a much larger aggregate burn rate.
This kind of user frustration is not unique to Anthropic; it echoes recurring debates across the AI industry about "quiet" changes to rate limits, quota calculations, or model routing that follow major model releases. Companies rarely publish granular token-cost comparisons between model versions, leaving power users to reverse-engineer behavioral shifts through anecdotal usage patterns, as seen here. For subscribers paying a flat monthly fee for tiered usage (like Max 5), an increase in per-task token cost effectively amounts to a stealth reduction in service, even if no explicit policy change was announced—a scenario that tends to generate significant community backlash when discovered by prosumer or developer audiences who rely on predictable throughput for professional workflows.
More broadly, this incident reflects the growing pains of the shift toward agentic AI systems, where usage is no longer measured in simple chat turns but in complex, multi-step, multi-agent orchestrations whose resource consumption is difficult for both users and providers to predict or communicate clearly. As frontier labs like Anthropic push models toward greater autonomy and longer-horizon task execution, the mismatch between flat-rate subscription pricing and highly variable, model-dependent compute costs is likely to become a recurring flashpoint. Expect continued scrutiny from power users on model upgrades, more demands for transparent token-accounting tools, and pressure on providers to either explain cost deltas explicitly or adjust plan limits in tandem with model releases to avoid appearing to silently throttle their most engaged customers.
Read original article →