← X
X

Each time we release a model, we run the same test: give it code that trains a s

X · AnthropicAI · 2026-06-04
Anthropic runs a standardized test with each model release by providing code to optimize AI model training speed, a task requiring skilled humans 4-8 hours to achieve 4x acceleration. Claude Opus 3 achieved approximately 3x speedup in May 2024, while Mythos Preview achieved approximately 52x speedup in April.

Detailed Analysis

Anthropic has disclosed a striking internal benchmark result that illustrates the accelerating capability of its AI systems: a standardized test that asks each newly released model to optimize code for training a small AI model has produced dramatically diverging results over approximately two years. Where a skilled human engineer requires four to eight hours to achieve a 4x performance speedup on this task, Claude Opus 4 in May 2024 averaged roughly a 3x speedup, falling just short of the human baseline. By April 2026, the company's Mythos Preview model achieved approximately a 52x speedup — more than seventeen times the gain produced by its predecessor and far exceeding what human engineers typically accomplish in the allotted time. The significance of this result lies not merely in the absolute number but in the shape of the progression. Observers in the social media discussion surrounding Anthropic's announcement noted that the trajectory — from near-human parity in mid-2024 to a result nearly an order of magnitude beyond human performance in 2026 — does not resemble a linear scaling trend. This non-linearity suggests that compound gains in model capability, evaluation tooling, and training infrastructure may be reinforcing one another in ways that produce phase-shift-like jumps rather than smooth incremental improvements. The benchmark itself is particularly revealing because it measures a task central to AI development: optimizing the very kind of code that produces AI systems, making strong performance on this test a potential indicator of how effectively a model could assist or accelerate the development of subsequent models. The broader strategic implications are substantial. Anthropic's benchmark measures a form of AI-assisted AI research, placing it directly at the intersection of what has long been speculated about as recursive self-improvement — the capacity of AI systems to meaningfully accelerate their own development pipeline. Separately reported figures from the thread suggest Mythos Preview has also outpaced human researchers on decision-accuracy tasks by 64%, with that margin reportedly tripling since the beginning of 2026, though those figures come from less verified sources within the discussion. What appears more substantiated is the code optimization result, which Anthropic is presenting as a consistent internal yardstick applied across model generations. The commentary from investors and analysts responding to the announcement reflects a debate about where competitive advantage will ultimately reside. Several voices in the discussion argue that the rapid capability gains shift the strategic moat away from proprietary model weights and toward proprietary training data and evaluation infrastructure — the "boring" compounding assets that make each successive model cheaper and faster to build. This framing positions Anthropic and its peers less as companies defined by a single flagship model and more as builders of flywheel systems in which each model generation produces the tooling and data that improves the next. One unverified and potentially fabricated claim in the thread alleges national security implications involving Mythos and classified government systems; that claim lacks credible sourcing and should be treated with significant skepticism. Taken together, the benchmark data Anthropic has released marks a notable public disclosure of the pace at which frontier AI systems are moving from approximate human parity on complex technical tasks to substantial superhuman performance. The code optimization test is a narrow but meaningful proxy for a category of capability — automated software engineering and systems optimization — that has direct commercial and scientific consequence. The speed of movement along that benchmark, from 3x to 52x in roughly two years, will likely intensify both investor attention and regulatory scrutiny around the degree to which AI systems can now meaningfully participate in accelerating their own development cycle.
Tweet screenshot
Read original article →

Don't Miss a Deploy

Claude moves fast. Get the signal — no noise — straight to your inbox every morning.