← Reddit

In one year, Claude went from being able to solve ~none of the hardest math problems to solving almost all of them

Reddit · EchoOfOppenheimer · June 13, 2026

Detailed Analysis

Claude's mathematical reasoning capabilities have undergone a dramatic transformation over approximately a twelve-month period, with Anthropic's flagship model progressing from near-zero performance on the most challenging mathematical benchmarks to achieving near-complete solve rates. The claim, circulating on Reddit with accompanying visual evidence, highlights a trajectory that would have seemed implausible just two years ago and underscores the accelerating pace of capability development in frontier AI systems. While the specific benchmark referenced is not named in the post, the description aligns closely with documented improvements on competition-level mathematics such as AIME (American Invitational Mathematics Examination) problems and similar olympiad-style challenges, where early Claude versions consistently struggled.

The gains are attributable in large part to architectural and training advances Anthropic introduced through successive model generations. The introduction of extended thinking capabilities in Claude 3.7 Sonnet marked a significant inflection point, allowing the model to engage in prolonged chain-of-thought reasoning before producing a final answer — a feature particularly well-suited to multi-step mathematical proofs and competition problems. Subsequent releases in the Claude 4 family continued this trajectory, with reinforcement learning techniques applied specifically to reasoning tasks further sharpening the model's ability to decompose complex problems, apply formal mathematical logic, and self-correct errors mid-computation. The compounding effect of these successive improvements helps explain the steep, near-vertical improvement curve that the referenced image appears to illustrate.

The significance of this development extends well beyond benchmark performance. Competition-level mathematics has long served as a proxy for deep, generalizable reasoning ability — problems at the AIME or olympiad level cannot be solved by pattern matching or retrieval alone; they require genuine inference under uncertainty, multi-step planning, and error recovery. A model that can reliably solve such problems is, in principle, capable of applying similar reasoning to scientific research, engineering design, and formal verification tasks. This is why mathematicians and AI researchers alike treat these benchmarks as meaningful signals rather than academic curiosities.

The one-year timeframe is itself a striking data point about the current pace of AI development. Progress that might historically have taken a decade — matching and then surpassing skilled human performance on elite mathematical competition problems — is now occurring within a single product cycle. Anthropic is not alone in this trajectory; OpenAI's o-series models and Google DeepMind's Gemini variants have followed comparable improvement curves on mathematical reasoning tasks, suggesting that the underlying drivers — scaled compute, improved training recipes for reasoning, and test-time compute strategies — are broadly applicable across frontier labs rather than idiosyncratic to any single approach.

Broader implications center on questions of trust, deployment, and what comes next. As Claude approaches saturation on existing hard-math benchmarks, the research community faces pressure to develop new evaluation frameworks that can continue to meaningfully differentiate model capabilities. Organizations like Epoch AI have introduced benchmarks such as FrontierMath specifically to stay ahead of this saturation problem, posing problems that require research-level mathematical creativity. The speed with which Claude has closed the gap on previously intractable problems signals that those harder benchmarks may themselves face obsolescence sooner than anticipated, reinforcing a broader pattern in which AI capabilities consistently outpace the evaluation infrastructure designed to measure them.

Article image Read original article →