← Anthropic News

More details on Fable 5’s cyber safeguards and our jailbreak framework

Anthropic News · July 3, 2026
Claude Fable 5 has been redeployed globally with cybersecurity safeguards featuring AI classifiers designed to detect and block dangerous or potentially dangerous cybersecurity uses. Anthropic introduced a proposed AI jailbreak severity framework developed with Glasswing partners to provide consistent terminology for describing how severely a jailbreak compromises model safety features. The framework establishes four categories of cybersecurity activities—prohibited use, high-risk dual use, low-risk dual use, and benign use—to address the dual-use nature of many cybersecurity capabilities.

Detailed Analysis

Anthropic's disclosure about Claude Fable 5's cybersecurity safeguards represents a notable step toward transparency in how the company manages one of the most technically fraught dual-use problems in AI safety. The post details a four-tier classification system—prohibited use, high-risk dual use, low-risk dual use, and benign use—that governs how Fable 5's safety classifiers respond to cybersecurity-related prompts. This taxonomy attempts to formalize a distinction that has long vexed AI safety teams: the same technical capability, such as scanning code for vulnerabilities or understanding malware obfuscation techniques, can serve legitimate defensive purposes or enable malicious attacks depending entirely on user intent, which is notoriously difficult for a classifier to infer from a single prompt.

The concept of a "safety margin," illustrated through a boundary diagram reproduced from Anthropic's earlier Fable redeployment post, is central to understanding how these classifiers actually operate in practice. Rather than attempting to perfectly separate benign from harmful requests—an arguably impossible task given the ambiguity inherent in dual-use technology—Anthropic deliberately widens the zone of blocked behavior to include some genuinely benign requests alongside low-risk dual-use ones. This produces a known tradeoff: higher false-positive rates, meaning legitimate security researchers or defenders occasionally get blocked, in exchange for greater confidence that genuinely dangerous jailbreak attempts get caught. The company explicitly states it set this margin larger for Fable 5 than for prior models, suggesting either increased capability in the underlying model (making misuse more consequential) or lessons learned from earlier deployments that prompted a more conservative posture.

Perhaps more significant than the classifier taxonomy itself is Anthropic's introduction of a draft "jailbreak severity framework," developed in partnership with an entity referred to as "Glasswing." This addresses a genuine gap in AI governance: there is currently no shared vocabulary or standard for describing how severe a given jailbreak is, which makes it difficult for AI labs, researchers, and governments to communicate consistently about risk. A jailbreak that unlocks a minor, low-stakes behavior is categorically different from one that strips away all cybersecurity safeguards simultaneously, yet without a common severity scale, incident reports and policy discussions risk talking past each other. By publishing an early draft and inviting feedback from academia, industry, civil society, and government—alongside launching a HackerOne bug-bounty-style program specifically for cyber jailbreaks—Anthropic is positioning itself as a convener trying to establish an industry-wide standard rather than a unilateral internal policy.

This move fits into a broader pattern across the frontier AI industry of formalizing risk communication as models become more capable in domains with direct security implications. As models like Claude increasingly demonstrate competence in code analysis, vulnerability discovery, and technical reasoning, the dual-use dilemma in cybersecurity becomes analogous to concerns long debated in biosecurity and chemical weapons contexts—capabilities that empower defenders also inherently empower attackers, and the gap between "helpful" and "harmful" is often a matter of context and intent rather than technical content. Anthropic's willingness to publish granular examples of what falls into each risk category, and to open its jailbreak severity framework to external critique before finalizing it, reflects a strategic bet that transparency and coalition-building around shared standards will serve the company better than either overly permissive deployment or so restrictive an approach that it undermines the model's usefulness to legitimate security professionals. Whether other major AI labs adopt or contest this proposed framework will be a meaningful signal of whether the industry is converging toward common risk-communication standards or continuing to develop fragmented, company-specific safety approaches.

Article image Article image Read original article →