← X
X

New Frontier Red Team blog: Phase 2 of Project Fetch, where we test how well Cla

X · AnthropicAI · 2026-06-18
The New Frontier Red Team tested Phase 2 of Project Fetch to evaluate Claude's ability to program a robodog. Opus 4.7 demonstrated approximately 20x faster performance compared to the previous year's best human team assisted by Opus 4.1, though the robodog ultimately failed to fetch a beach ball.

Detailed Analysis

The New Frontier Red Team's Phase 2 results from Project Fetch represent a striking benchmark in autonomous AI coding capability applied to robotics. The core finding — that Claude Opus 4.7, operating independently, completed its programming work approximately 20 times faster than the prior year's top human team working in collaboration with Claude Opus 4.1 — marks a dramatic generational leap in raw throughput between successive model versions. The project tasks AI systems with writing the software necessary to control a quadrupedal robot, commonly referred to as a robodog, in order to retrieve a beach ball, a deceptively complex challenge that requires integrating perception, locomotion planning, and object manipulation logic. The comparison embedded in this result is particularly telling. The 2025 benchmark involved skilled human programmers augmented by Claude 4.1, representing the best of human-AI collaborative coding at that time. In 2026, Claude Opus 4.7 replaces that entire human-AI team and outpaces it by a factor of twenty in speed. This reflects not merely incremental model improvement but a qualitative shift in the viability of fully autonomous AI agents for complex, multi-step technical tasks. The framing of a "red team" conducting these tests also signals that this is rigorous, adversarial evaluation rather than a controlled demonstration, lending greater credibility to the speed differential reported. Critically, however, speed did not translate into success. The robodog still failed to fetch the beach ball, a result that underscores a persistent and important gap in AI-driven robotics: the ability to generate code rapidly is categorically distinct from the ability to generate code that correctly solves a physically embodied problem. Robotics programming involves tight feedback loops between software behavior and real-world physics, sensor noise, and mechanical constraints — domains where sheer coding velocity provides limited advantage if the underlying logic or model of the environment remains flawed. This result connects to a broader tension emerging in 2026-era AI development: frontier models are demonstrating superhuman performance on coding benchmarks and agentic software tasks, yet consistently struggle when those software outputs must interface with the physical world. The gap between language-model reasoning and reliable robotic task completion remains one of the most consequential unsolved problems in applied AI. Projects like Fetch serve as important empirical probes of exactly where that gap lies, revealing that autonomy and speed in code generation have scaled rapidly while real-world execution fidelity has not kept pace. The trajectory suggested by comparing Phase 2 to Phase 1 also implies that future iterations may close the execution gap as models continue to improve. A 20x speed advantage over human teams means more iterations can be tested per unit time, potentially accelerating the feedback loop through which robotic control code is debugged and refined. Whether that acceleration will be sufficient to bridge the physical-world reliability gap — or whether fundamentally different architectures for embodied AI will be required — remains the central question Project Fetch appears designed to answer.
Tweet screenshot
Read original article →

Don't Miss a Deploy

Claude moves fast. Get the signal — no noise — straight to your inbox every morning.