← Reddit

tested Claude Fable 5 and Opus 4.8 across 917 coding-agent scenarios. Fable won by 0.9 points.

Reddit · rohansrma1 · June 12, 2026
A comparison of Claude Fable 5 and Opus 4.8 across 917 coding-agent scenarios found Fable 5 achieved a 92.9 overall score compared to Opus 4.8's 92.0 score, representing a 0.9-point improvement. However, Fable 5 costs approximately 73% more at $1.25 per task versus Opus 4.8's $0.74 per task. Fable 5 refused 26 tasks that Opus completed successfully, with failures including security-review and bioinformatics tasks.

Detailed Analysis

A Tessl-authored benchmark pitting Claude Fable 5 against Opus 4.8 across 917 coding-agent scenarios reveals a competitive but narrow performance gap between what Anthropic has branded its first public Mythos-class model and its predecessor. Fable 5 achieved an overall score of 92.9 compared to Opus 4.8's 92.0 — a margin of just 0.9 points — while costing approximately $1.25 per task versus $0.74 for Opus 4.8. That translates to a roughly 73% cost premium for a sub-one-point improvement on the specific workloads tested, a disparity that immediately frames the core practical question the article raises: whether the capability uplift is sufficient to justify the added expenditure at scale. The tasks were drawn from skills listed in Tessl's own registry and scored using its internal task evaluation framework, which the author discloses, allowing independent measurement of both model and skill contributions to performance.

One of the more substantive findings is that Fable 5 outright refused 26 tasks that Opus 4.8 completed successfully. These refusals spanned security-review workflows and routine bioinformatics pipelines — neither of which would typically be considered high-risk territory. Anthropic has reportedly acknowledged that the initial rollout of Fable 5 was configured with overly conservative safety guardrails, suggesting the refusal rate is a temporary calibration issue rather than a permanent architectural constraint. Still, for production coding-agent deployments where reliability and task completion rates are critical, even a transient over-restriction poses real operational friction, particularly in specialized domains like bioinformatics where the model's hesitancy appears to have been miscalibrated relative to actual risk.

The benchmark carries meaningful caveats that contextualize its conclusions. Because the evaluation tasks originate from Tessl's own skill registry and are scored by Tessl's internal framework, the results reflect performance on a specific, curated distribution of coding-agent scenarios rather than a universally representative sample. The author transparently discloses a conflict of interest as a Tessl employee, and the methodology links to public documentation, which allows for some degree of independent scrutiny. Nevertheless, the findings are most reliably generalized to workflows that resemble Tessl's registry — agentic coding tasks with structured inputs and measurable outcomes — rather than to the full breadth of use cases for which either model might be deployed.

In the broader context of AI development, this comparison illustrates a recurring dynamic in frontier model releases: successive generations deliver incremental performance gains that are statistically real but practically modest on established benchmarks, while introducing new tradeoffs — in this case, higher cost and initially tighter safety restrictions. The introduction of the "Mythos-class" branding suggests Anthropic is deliberately stratifying its model lineup by capability tier and target use case, a pattern mirrored across the industry as providers attempt to match pricing to workload complexity. The Fable 5 vs. Opus 4.8 gap also reinforces a growing observation in the field that for well-defined, structured agentic tasks, older high-performing models remain highly competitive against newer releases, and that cost-per-task economics are becoming as decisive a selection criterion as raw benchmark scores for engineering teams operating at scale.

Article image Read original article →