← Reddit

Why are Sonnet 5 and Opus 4.8 so sensitive about Security all of a sudden?

Reddit · secpoc · July 2, 2026
A user reported experiencing suspended and paused conversations when using Sonnet 5 and Opus 4.8 models for security-related inquiries, including questions about CVE issuance processes and drafting emails to cybersecurity colleagues. The suspensions rendered the models practically unusable over several days, prompting the user to revert to Claude Opus 4.6.

Detailed Analysis

A Reddit thread in r/Anthropic has surfaced user complaints that Claude's newer models—Sonnet 5 and Opus 4.8—are exhibiting unusually aggressive conversation-suspension behavior around security-related topics. The original poster describes having a chat paused simply for asking about the CVE (Common Vulnerabilities and Exposures) issuance process, a standard and publicly documented mechanism used throughout the cybersecurity industry to catalog software vulnerabilities. In a separate instance, drafting a routine email to a colleague in a cybersecurity department also triggered a suspension. Frustrated by the pattern, the user reports reverting to the older Opus 4.6 model, noting they are enrolled in Anthropic's Claude for Very Important People (CVP) or similar early-access program, which presumably gives them some flexibility in model selection.

This complaint sits within a well-documented tension in AI safety engineering: the tradeoff between robust misuse prevention and false-positive over-triggering on benign requests. CVE numbers, vulnerability disclosure processes, and cybersecurity terminology are foundational knowledge for millions of IT professionals, researchers, and students, and they are widely available on public resources like MITRE's CVE database and NIST's NVD. When a model's classifiers flag legitimate professional inquiries about these topics as potentially malicious—perhaps conflating "vulnerability" or "exploit" terminology with attempted jailbreaks or requests for offensive tooling—it signals that safety classifiers may have been tuned too conservatively, likely in response to broader concerns about AI models being used to accelerate cyberattacks or generate malware.

The timing described in the post, with the behavior emerging "over the past few days" following what appears to be a model or safety-system update, suggests this may be tied to a specific classifier or system-prompt change rather than a fundamental limitation of the underlying models. Anthropic has historically iterated rapidly on its "Constitutional AI" and usage-policy enforcement mechanisms, and such systems are known to sometimes overcorrect after being retrained or retuned—especially following incidents elsewhere in the industry involving AI-assisted cyber threats. Enterprise and security-industry customers are particularly sensitive to this kind of friction because cybersecurity professionals are a core Claude user base, relying on the models for tasks like vulnerability triage, security documentation, and threat analysis. If legitimate security work is being mistaken for malicious intent, it directly undermines Claude's utility for exactly the audience most invested in secure software practices.

More broadly, this incident reflects a recurring pattern across the AI industry: as models grow more capable, companies tighten guardrails around dual-use domains like cybersecurity, biology, and chemistry, and these tightenings frequently produce collateral damage against benign professional use cases before being recalibrated. The episode also highlights how quickly user sentiment can shift in response to invisible backend changes—Anthropic did not, per the post, publicly announce a change in security-topic handling, leaving users to diagnose the shift through trial and error and informal community discussion on forums like Reddit. Such friction points are likely to keep surfacing as Anthropic balances safety commitments (including its public pledges around responsible scaling and misuse prevention) against the practical needs of professional users who depend on frontier models for legitimate security work, and the CVP user's fallback to an older, less restrictive model version underscores how regressions in usability can quickly erode trust in newer releases even when they offer other capability improvements.

Read original article →