← X

Correction: Claude Opus 4's ~3x average speedup dates to May 2025, not May 2024.

X · AnthropicAI · June 4, 2026
Anthropic issued a correction regarding Claude Opus 4's speedup metric, clarifying that the approximately 3x average speedup dates to May 2025 rather than May 2024. The evaluation itself only existed since September 2024, but when backtested on earlier models from May 2024, no speedup was detected.

Detailed Analysis

Anthropic issued a factual correction clarifying that Claude Opus 4's approximately three-times average speedup in a research decision-making evaluation benchmark dates to May 2025, not May 2024 as previously stated. The distinction is significant: the evaluation framework itself only came into existence in September 2024, making the earlier date technically impossible to verify through live testing. When Anthropic backtested the benchmark against models from May 2024, those older systems showed no measurable speedup at all, underscoring that the capability improvement is genuinely recent and represents a discrete jump rather than a gradual progression.

The correction surfaces in the context of broader, more sweeping claims circulating in public discourse about AI research capability acceleration. Social media commentary appended to the notice references a "Mythos Preview" evaluation purportedly showing AI decision-making accuracy exceeding human researchers by 64%, as well as assertions that a performance curve has moved from 0x to 3x to 52x in under three years. These figures, while striking, lack sourced verification in the material provided, and at least one post attributes dramatic national-security implications to an AI system called "Mythos" in a manner that reads as speculative or unverified. The correction itself, by contrast, reflects Anthropic's practice of issuing precise retractions when benchmark data is mischaracterized — a form of methodological transparency that is comparatively rare in AI industry communications.

The underlying measurement infrastructure is itself notable. The existence of a formal, repeatable evaluation framework for AI research decision speed — one sophisticated enough to allow meaningful backtesting across model generations — represents a maturation in how Anthropic approaches capability assessment. Prior to September 2024, no such instrument existed, meaning earlier models simply could not be evaluated on this dimension in any contemporaneous sense. The ability to retrospectively apply new benchmarks to older models, and to find a clean null result for May 2024 systems, gives the May 2025 improvement additional credibility as a genuine inflection rather than an artifact of measurement design.

The broader commentary surrounding the correction reflects a recurring tension in AI discourse between careful empirical claims and accelerationist narratives. Observers pointing to exponential curves and "phase changes" are responding to real data points — Claude Opus 4's 3x speedup is a documented, if newly corrected, finding — but the leap from a controlled research benchmark to claims about recursive self-improvement or national-security breaches involves substantial inferential distance. Anthropic's willingness to issue a correction on a relatively minor dating error stands in contrast to the more hyperbolic framing that surrounds the same underlying performance data in public commentary.

For the AI development field more broadly, the episode illustrates how benchmark provenance and timeline accuracy increasingly matter as performance claims become consequential for policy, investment, and public trust. A three-times improvement in AI research decision-making speed, properly dated to a ten-month window in 2024–2025, is a meaningful signal about capability trajectory. Misattributing it to an earlier date — even by one year — distorts the apparent rate of progress and could mislead downstream analysis. Anthropic's correction, while modest in scope, reinforces that rigorous documentation of when capabilities emerge is as important as documenting what those capabilities are.

Tweet screenshot Read original article →