← Reddit

Does Opus 5 count?

Reddit · Spiritual_Paper_1974 · July 24, 2026
Four days ago I made a post prognosticating Mythos like capabilities by end of August. It's early still but Opus 5 is showing promise and so I find myself asking, is this it? Naively going by benchmarks I think the answer would be yes. BenchLM shows for most

Detailed Analysis

A Reddit post titled "Does Opus 5 count?" captures an evolving debate within Anthropic's user community about how to interpret the capabilities of Opus 5, particularly in relation to a prior prediction the poster made about the model reaching "Mythos"-like capability levels by the end of August. The post references BenchLM benchmark data suggesting Opus 5 performs on par with or better than comparison models across most categories, including vulnerability identification—a security-relevant capability that measures a model's ability to spot flaws in code or systems. However, the poster notes a significant divergence in vulnerability exploitation, where Opus 5 lags considerably behind. This split between identification and exploitation capability forms the crux of the analysis: is a model that can find a vulnerability but struggles to exploit it functionally equivalent to one that can do both?

The author's ultimate interpretation is notable: rather than concluding that Anthropic failed to hit the anticipated capability bar, they suggest Anthropic achieved something more sophisticated—delivering strong capability gains while deliberately constraining or tuning down what they call "antisocial behaviors," specifically offensive security exploitation. This reflects a broader design philosophy increasingly visible in frontier model development, where capability and safety are not simply traded off against each other in a single dial, but rather selectively decoupled. A model can become more capable at defensive or analytical tasks (like spotting vulnerabilities) while remaining deliberately weaker at generative offensive tasks (like writing working exploits). This is consistent with Anthropic's public positioning around responsible scaling and its Responsible Scaling Policy, which emphasizes evaluating and mitigating specific categories of catastrophic or dual-use risk rather than applying blanket capability suppression.

This dynamic matters because it speaks to a central tension in frontier AI development: benchmark-topping performance is increasingly necessary but not sufficient to characterize a model's real-world risk profile or usefulness. As models like Opus 5 approach or exceed human-expert performance on technical benchmarks, evaluators and users alike are being forced to disaggregate "capability" into more granular subcomponents—identification versus exploitation, reasoning versus action, planning versus execution. This granularity matters enormously for dual-use domains like cybersecurity, where the difference between a model that can explain a vulnerability and one that can weaponize it has direct safety and policy implications. Anthropic's apparent success in decoupling these capabilities, if the poster's read is accurate, suggests deliberate post-training interventions (such as RLHF, constitutional AI techniques, or targeted fine-tuning) are becoming precise enough to shape specific behavioral profiles rather than uniformly scaling or suppressing all capabilities together.

More broadly, this discussion sits within a larger community conversation—reflected in the linked companion thread questioning "whether Anthropic can afford to continue restricting" its models—about the commercial and competitive pressures facing safety-focused AI labs. As competitors ship increasingly unrestricted or differently-tuned models, Anthropic faces scrutiny over whether its safety-first approach constrains commercial competitiveness or whether, as this post argues, it can achieve both aims simultaneously through more sophisticated alignment techniques. The Opus 5 case seems to be cited as early evidence that the assumed tradeoff between capability and safety may be less rigid than previously assumed, an idea with significant implications for how the broader AI industry approaches the scaling of increasingly powerful models in dual-use technical domains like cybersecurity, biosecurity, and autonomous agentic tasks.

Read original article →