← Reddit

Fable passes the "When A.I. Passes This Test, Look Out" test

Reddit · droidment · June 12, 2026
Claude Fable achieved a 53% score on a benchmark test designed to measure AI systems' capability to answer questions accurately across any topic. The result came approximately six months later than expert predictions that anticipated AI systems would reach this performance threshold by the end of 2025. At this level of performance, AI systems are considered capable of functioning as "world-class oracles," answering questions more accurately than human experts.

Detailed Analysis

Anthropic's Claude Fable model has crossed the 53% threshold on Humanity's Last Exam (HLE), a benchmark widely regarded as one of the most rigorous tests of AI capability ever constructed. The milestone was flagged in a January 2025 New York Times article featuring Dan Hendrycks, a prominent AI safety researcher, who described the exam as a collection of extraordinarily difficult questions spanning mathematics, science, humanities, and other domains — questions designed specifically to stump frontier AI systems. Hendrycks predicted at the time that AI scores on HLE would surpass 50% by the end of 2025, a threshold he characterized as the point at which AI systems could be considered "world-class oracles," capable of outperforming human experts across virtually any topic.

The achievement arrived approximately six months behind Hendrycks's original timeline, placing the milestone in mid-2026 rather than late 2025. This delay is itself analytically significant. While the broader trajectory of AI capability improvement matched Hendrycks's directional forecast, the pace was modestly slower than anticipated — a pattern consistent with the general difficulty of predicting exact inflection points in AI development, even for researchers with deep domain expertise. The six-month slip does not undermine the underlying prediction so much as illustrate the inherent imprecision of capability forecasting in a field where progress is nonlinear and model architectures evolve through discrete, sometimes irregular release cycles.

The framing of HLE as a threshold test — rather than a gradual performance metric — carries important implications for how this result is likely to be interpreted by the research community and the public. By defining 50% as the point at which AI transitions into "world-class oracle" territory, Hendrycks effectively created a binary narrative: before and after. Claude Fable's 53% score pushes the technology past that narrative boundary, lending rhetorical weight to claims that AI has entered a qualitatively new phase of capability. Whether that framing accurately captures the practical implications remains contested, but it shapes public and institutional perception in ways that go beyond what the raw benchmark number alone would convey.

In the broader context of AI development, this result reflects Anthropic's continued competitive positioning at the frontier of large language model capability. Claude Fable's performance on HLE adds to a growing body of evidence — across reasoning benchmarks, coding evaluations, and scientific problem-solving tasks — that the most capable models are now operating in domains previously considered exclusive to human specialists. The broader industry trend suggests that benchmark saturation is accelerating: tests that were considered nearly impossible for AI systems only a few years ago are now being approached or surpassed within months of a benchmark's public release. HLE was specifically designed to resist this saturation, making Claude Fable's crossing of the 50% threshold a more durable signal than scores on older, more frequently trained-against evaluations.

The social and institutional stakes of this milestone extend well beyond the technical community. Hendrycks's "world-class oracle" framing, amplified by mainstream coverage in the New York Times, has primed a broad audience to treat the 50% threshold as a societal inflection point rather than a technical curiosity. As AI systems demonstrate the ability to answer expert-level questions more reliably than the experts themselves, questions about epistemological authority, professional gatekeeping, and the role of human expertise in high-stakes domains — medicine, law, scientific research, policy — become considerably more urgent. The fact that this threshold has now been crossed, even six months behind schedule, means those conversations are no longer purely speculative.

Read original article →