← Reddit

Sol ultra is dumber than Fable medium

Reddit · Available_Status1 · July 12, 2026
A user reported finding Sol ultra less capable than Fable medium, noting that Sol struggled with project comprehension and appeared to regurgitate information without understanding. The user questioned whether positive reviews praising Sol were genuine or the result of astroturfing, suggesting Sol performs adequately for standard programming but lacks sophistication for more complex tasks.

Detailed Analysis

A Reddit post in r/Anthropic titled "Sol ultra is dumber than Fable medium" captures a user's frustrations after comparing OpenAI's ChatGPT "Sol" model against Anthropic's "Fable" model (evidently code names or community shorthand references, possibly to obfuscated or renamed versions of GPT and Claude models respectively, given the subreddit context and comparison to Opus 4.8). The poster describes returning to ChatGPT after having abandoned it previously, only to find that "Sol medium" underperformed relative to Opus 4.8, misunderstanding basic project context and seemingly regurgitating prior outputs from "Fable" without genuine reasoning. The user questions whether the wave of positive sentiment around Sol is organic or the product of coordinated astroturfing, a suspicion increasingly common in AI discourse where corporate reputations and model rankings are contested in public forums.

The post is notable less for hard technical findings and more as a data point in the ongoing, highly subjective battle for perceived AI supremacy between Anthropic's Claude models and OpenAI's offerings. Community-driven comparisons like this one are a significant, if unscientific, feedback mechanism shaping public perception of model quality. Users frequently rely on anecdotal side-by-side testing—especially around coding, reasoning, and project comprehension—to decide which assistant to trust for professional or technical work. The reference to triggering "safety" behavior via a "/code-review" command, which allegedly degrades output quality ("then fable is a brick too"), also touches on a persistent tension in the AI industry: the trade-off between safety guardrails and raw model capability. Users often perceive safety-triggered response modes as a tax on usefulness, a friction point Anthropic and competitors alike continually navigate as they tune models like Claude (marketed for its strong safety posture) against rivals optimized more aggressively for raw performance benchmarks.

This kind of post also underscores the challenge of exact-model comparison in an era of rapid, iterative releases. With Anthropic pushing versions like Opus 4.8 and OpenAI iterating through its own family of models, users are often left comparing tools across different pricing tiers ("ultra," "medium") and unclear naming conventions, making objective quality assessment difficult. The poster's closing acknowledgment—that the discrepancy might stem from prompting style rather than inherent model capability—reflects a broader, underappreciated truth in AI usage: perceived intelligence gaps between top-tier LLMs are often as much a function of user technique, project setup, and context management as they are of underlying model architecture or training.

Broadly, this Reddit thread is emblematic of the increasingly tribal, comparison-heavy culture surrounding frontier AI models, where communities dedicated to specific labs (like r/Anthropic) serve as informal battlegrounds for reputation management, product loyalty, and crowd-sourced benchmarking. As both Anthropic and OpenAI continue to release incremental updates and adjust safety mechanisms, this kind of qualitative, user-driven discourse will likely remain a key—if messy—signal in the broader narrative of which AI lab is "winning" the intelligence race, even as it's often confounded by naming ambiguity, subjective prompting differences, and the ever-present specter of manufactured hype.

Read original article →