← Reddit

My profile prompts to fix Opus 5's biggest complaints (timid, unfocused, doesn't trust itself)

Reddit · EverGreenMob · August 1, 2026
A user published custom instructions designed to counter Opus 5's perceived weaknesses—timidity, task abandonment under pushback, and unsolicited tangents—through three core rules: defaulting to recommendations over extended reasoning, arguing once before yielding to user preferences, and staying focused on the user's request without unsolicited analysis. The instructions establish a warm, humble persona that validates correct thinking before offering corrections, with emphasis on directive guidance, graceful disagreement, and respecting user preferences while considering overall wellbeing.

Detailed Analysis

Reddit users are already engineering workarounds for behavioral quirks in Claude Opus 5 just as early reviews of the model surface a consistent set of complaints: excessive timidity, a tendency to abandon tasks at the first sign of user pushback, unsolicited tangents and pop-psychology commentary on the user's state of mind, and a curious reluctance to assert its own judgment even when directly asked. One reviewer's anecdote — asking the model point-blank who was smarter, her or it, and getting a non-answer — has become a shorthand for the broader critique that Opus 5, despite gains in measured alignment, feels evasive and self-doubting in everyday use.

The apparent contradiction is notable: Anthropic's own system card reportedly documents Opus 5 as the company's most aligned release yet, with the lowest rates of deceptive behavior among recent Claude models. That framing suggests the friction users are experiencing isn't a refusals problem in the traditional sense — the model isn't stonewalling requests — but rather a personality-tuning problem, where safety-oriented caution has bled into general conversational confidence. This distinction matters because it reframes "alignment" as a multidimensional target: a model can score well on deception and harm metrics while still failing users on responsiveness, decisiveness, and staying on-topic. It's a reminder that benchmark-level alignment claims don't automatically translate into a better subjective user experience, and that qualitative complaints from power users often surface issues that formal evaluations don't capture.

The community response — sharing custom system prompts built around rules like "default to making the call," "make your case once, then yield," and "stay focused on the ask" — illustrates how much of the burden for shaping model behavior has shifted onto end users through prompt engineering rather than being solved at the model or platform level. The fact that these homemade personas are explicitly designed to counteract launch-day complaints suggests a recurring pattern in frontier model releases: initial versions ship with a default persona tuned by the vendor for broad safety and inoffensiveness, and the power-user community rapidly iterates on custom instructions to recover the assertiveness and focus that got dialed down. The reference to a private benchmark comparing Opus 5 unfavorably to a competing "Fable 5" model on a "gentle pushback" test — while also noting that custom instructions can close the gap — reinforces that persona-level tuning is increasingly treated as a solvable, almost cosmetic layer on top of the underlying model capability.

More broadly, this episode reflects a maturing tension in commercial LLM deployment: as companies like Anthropic push harder on measurable alignment and safety metrics ahead of major releases, they risk overcorrecting into behaviors — hedging, disclaiming, moralizing asides, non-committal answers — that erode trust and usability for sophisticated users who want a confident collaborator rather than a cautious assistant. The rapid emergence of detailed, publicly shared system prompts to "fix" these tendencies also signals how central prompt customization has become to the Claude ecosystem, effectively turning users into informal product designers who patch perceived deficiencies faster than official updates can. As frontier labs compete not just on raw capability but on the felt quality of interaction, expect persona tuning, default system prompts, and user-configurable "directives" to become as consequential a battleground as benchmark scores themselves.

Read original article →