← Reddit

I switched from Claude to Codex the day GPT-5.5 launched, and the gap isn't subtle — it's Apple vs. Microsoft, circa 1998.

Reddit · Inevitable_Raccoon_9 · July 11, 2026
On Codex, the fundamentals just work. Multiple chats in multiple windows: open and go. Start a conversation on my Mac, finish it on my phone mid-thought — seamless, zero friction, no workaround required. It behaves the way software in 2026 is supposed to

Detailed Analysis

A Reddit post comparing Anthropic's Claude desktop app unfavorably to OpenAI's Codex has surfaced pointed criticism about the gap between model quality and product execution at Anthropic. The author's central claim is not that Claude's underlying models are inferior — they explicitly note the two are "comparably capable" — but that the surrounding software experience is significantly worse. Specific complaints include poor multi-window support, lack of seamless cross-device handoff (starting a conversation on desktop and continuing on mobile), and a general sense that the desktop app feels unfinished or unstable. The framing invokes a dated but evocative analogy: Anthropic's desktop client behaving like "Windows circa 1998," while Codex behaves like a modern, polished piece of software users expect in 2026.

This kind of critique matters because it highlights a recurring tension in AI company strategy: research-and-model excellence versus product-and-engineering execution. Anthropic has built its reputation primarily on model capability, safety research, and enterprise/developer trust — areas where it has often been seen as at or near the frontier alongside OpenAI and Google DeepMind. However, model quality alone doesn't guarantee user retention, especially as competitors ship increasingly capable models of their own. When multiple labs converge on similar levels of reasoning, coding, and conversational ability, the differentiating factor increasingly becomes the day-to-day usability of the product wrapped around the model — session continuity, interface responsiveness, multi-platform sync, and general polish. If users perceive parity in model quality, friction in the product becomes the deciding factor in retention and willingness to pay a premium subscription.

This tension is especially relevant to Anthropic given its dual go-to-market approach: consumer-facing apps (Claude.ai, desktop, mobile) alongside a heavy developer/enterprise focus via the API and tools like Claude Code. Anthropic has historically prioritized safety, alignment research, and enterprise API robustness over consumer app polish, and this critique suggests that trade-off may be catching up with them as competitors like OpenAI move quickly to close capability gaps while maintaining strong product execution across ChatGPT, Codex, and related surfaces. The launch of GPT-5.5 evidently intensified this comparison — when a rival ships a comparably strong model without sacrificing UX, the incumbent's product shortcomings become more visible and less excusable to paying customers.

More broadly, this reflects an industry-wide shift in AI competition: the "model war" is increasingly becoming a "product war." As foundation models across major labs converge toward similar levels of raw capability — driven by scaling, distillation, and shared research advances — user experience, integration depth, platform reliability, and ecosystem completeness are emerging as the new competitive battlegrounds. Anthropic's challenge going forward will be proving that its safety-first, research-heavy culture can coexist with the kind of consumer product rigor that OpenAI, and to a lesser extent Google, have invested heavily in. Failure to close that gap risks eroding Anthropic's paid subscriber base even if its models remain technically excellent, reinforcing the broader lesson that in AI, shipping a great model is necessary but no longer sufficient to win users.

Read original article →