← Reddit

When should I use opus 4.8 vs Sonnet 5? I have Claude control my browser and fill a google chrome form which has lots of areas of data entry.

Reddit · TeaSipper007 · July 14, 2026
A user described frequent performance issues with Claude Opus 4.8 when automating Google Chrome form filling and data entry tasks, reporting instances of incorrect button clicks and requests to repeat instructions despite detailed documentation. The user provides the model with Excel files containing data for automatic form population but has experienced declining reliability.

Detailed Analysis

A Reddit user's query about choosing between Opus 4.8 and Sonnet 5 for a browser-automation task highlights a practical, increasingly common dilemma among Claude power users: how to match model selection to workload rather than defaulting to the most capable (and expensive) option. The poster describes a repetitive, structured task—reading data from an Excel file and using it to fill in a multi-field Google Chrome form—and notes that Opus, despite detailed instructions, has begun exhibiting "faffy" behavior: repeated corrections, misclicks, and inconsistent adherence to documented steps. This is a notable anecdotal data point because Opus-class models are generally marketed as Anthropic's top-tier reasoning models, intended for complex, ambiguous, or high-stakes tasks, while Sonnet-class models are positioned as faster, more cost-efficient, and often better suited to well-defined, high-volume operations.

The core issue reflects a broader pattern in agentic AI workflows: model reasoning depth and task determinism don't always align well. Tasks like form-filling from structured data are largely procedural—match column to field, click, type, move on—and don't require the kind of deep multi-step reasoning that Opus is optimized for. Ironically, more "thoughtful" models can sometimes overthink simple repetitive tasks, introducing unnecessary variation, hesitation, or self-correction loops that manifest as errors or inconsistency. Sonnet models, by contrast, are often tuned for speed and consistency on narrower, well-specified tasks, which may make them more reliable for exactly this kind of use case, even if their general reasoning ceiling is lower.

This also touches on the practical reality of Claude's role in computer-use and browser-automation contexts, where Anthropic has been expanding capabilities (via Claude's "Computer Use" API and integrations with browser control) to let the model interact directly with UI elements. In these settings, reliability of action execution—correctly identifying and clicking the right button, correctly mapping data to fields—is often more important than the depth of reasoning behind the action. Users operating in this space are increasingly finding that model choice isn't simply "bigger is better," but rather a matter of matching the model's strengths to the granularity and repeatability of the task at hand.

More broadly, this exchange reflects a maturing phase in how users engage with Claude's model lineup. Early adoption often defaults users toward the most powerful available model regardless of necessity, but as usage scales and workflows become more automated and repetitive, cost, speed, and consistency become larger factors than raw capability. Anthropic's multi-tier model strategy (Haiku, Sonnet, Opus) is designed precisely to let users make these tradeoffs, and community discussions like this one serve as informal benchmarking grounds where practitioners share real-world performance differences that don't always show up in official documentation or benchmarks. This kind of grassroots evaluation—comparing model behavior on concrete, repeatable business tasks—is increasingly valuable as agentic AI use cases proliferate beyond conversational assistance into direct software and browser control.

Read original article →