Detailed Analysis
A YouTube creator's hands-on comparison pits a purported new OpenAI release—referred to throughout the video by the internal-sounding names "GPT-5.6 Soul," "Terra," and "Luna"—against "Fable 5," evidently a codename for an unreleased or beta Claude model running through Claude Code's interface. Rather than leading with published benchmarks, which the creator dismisses as "boring," the test centers on two open-ended agentic coding tasks: building a playable browser-based bike game and generating an "interactive scroll-stopping website" from a single creative prompt. This kind of unstructured, vibe-based evaluation has become increasingly common in AI commentary as viewers grow skeptical of benchmark scores that don't always translate to perceived quality in real usage.
The results highlight a recurring tension in frontier model development: token efficiency versus output quality. In the bike-game test, the GPT-5.6-branded model completed its build faster and far more cheaply—about $4.50 versus $14.22 for the Claude-based model—while producing roughly a third of the output tokens (31,000 versus 90,000). Yet the creator judged the more expensive, more verbose Claude output to be substantially better, describing it as having a more explorable, GTA-like open-world feel compared to the more constrained top-down alternative. This illustrates a persistent theme in comparisons between OpenAI's Codex-style models and Anthropic's Claude: OpenAI's models tend to be leaner and cheaper per task, while Claude's models often generate richer, more elaborate outputs at higher token cost—a tradeoff that matters enormously for developers and businesses deciding which model to deploy for creative or front-end-heavy work versus high-volume, cost-sensitive applications.
The broader significance here lies less in the specific benchmark claims—about which independent confirmation is limited given the informal, creator-sourced nature of the names "Soul," "Terra," and "Luna"—and more in what the comparison reveals about how the AI industry is being evaluated in practice. Coding and generative-design tasks like game creation and interactive websites have become de facto stress tests for agentic coding ability, since they require sustained multi-step reasoning, creative judgment, and the integration of physics, visuals, and sound rather than single-shot code generation. As both OpenAI and Anthropic push iterative model updates at a rapid pace, independent creators testing models side by side—rather than relying solely on vendor-published benchmarks—have become an important, if unofficial, channel through which the market forms opinions about model quality, cost tradeoffs, and practical utility.
This dynamic also underscores the intensifying competitive pressure between Anthropic and OpenAI in agentic and coding-oriented use cases, an arena where Claude has built a strong reputation via Claude Code and where OpenAI is clearly positioning newer model variants to compete aggressively on price. Even if formal benchmarks show one model "winning" numerically, real-world creative and technical quality—as judged by actual builders—can diverge sharply, reinforcing that cost-per-token efficiency and raw benchmark performance are only part of the picture. As frontier labs continue to release faster, cheaper models, this kind of qualitative, task-based testing is likely to remain a key way both developers and casual users decide which assistant genuinely fits their workflow.
Read original article →