← Reddit

Dear Anthropic! Please Hear me Out! How am i supposed to use F5 with Current Limitations as an Ai Agent Security Project?

Reddit · RCBANG · July 2, 2026
The founder of Sunglasses.dev, a sandboxed AI agent security research project accepted into Anthropic's Cyber Verification Program, reported that Fable 5 automatically downgrades to Opus 4.8 when conversations contain security-related keywords such as "patterns," "false positive," or "prompt injections." Despite providing Anthropic with terminal session access for verification and conducting only sandboxed research activities, the model remained restricted, which the founder argued prevented legitimate security work needed to improve the project. The founder requested that Anthropic adjust its restrictions to enable established security projects to use Fable 5 within the temporary access period through July 7.

Detailed Analysis

A Reddit post from a developer identified only as "AZ," founder of Sunglasses.Dev, an AI agent security project, illustrates a friction point that frequently emerges when Anthropic rolls out new frontier models with cautious safety guardrails. The poster describes attempting to use a model referred to as "Fable 5"—apparently a codename or informal nickname circulating among early testers for a new Claude release, possibly Claude Opus 4.5 or a similarly numbered update—only to have the system automatically downgrade the conversation to an older model, Opus 4.8, mid-session. According to the author, this happened because the conversation contained cybersecurity-related terminology such as "patterns," "false positive," and "prompt injections," triggering an automated safety classifier despite the underlying work being defensive security research conducted in a sandboxed Docker environment rather than any live offensive capability.

The core tension here reflects a well-documented challenge in deploying increasingly capable AI models: balancing broad usability for legitimate technical and security work against the risk of misuse by bad actors. Anthropic, like other frontier AI labs, has invested heavily in usage policies and automated content classifiers designed to detect and restrict potentially harmful use cases involving hacking, malware, or offensive cyber operations. These classifiers often rely on keyword and pattern detection, which can produce false positives when legitimate security researchers use the same vocabulary as malicious actors. The poster's frustration is compounded by claims of having been accepted into Anthropic's Cyber Verification Program (CVP), a vetting initiative apparently intended to identify and support trusted security researchers precisely so they can work with Claude on sensitive security topics without triggering these restrictions. If accurate, this suggests the enforcement layer (automated downgrade/refusal systems) is not yet fully integrated with the trust layer (CVP vetting status), producing a disconnect where verified, sanctioned researchers still get blocked by the same guardrails meant to stop unverified bad actors.

This case is emblematic of a broader pattern in AI governance: as safety systems become more aggressive and models more capable, the community of "trusted insiders"—security researchers, red-teamers, penetration testers, and academic labs—increasingly finds itself caught in the crossfire of restrictions designed for anonymous or malicious users. Anthropic has publicly emphasized cybersecurity as both an area of AI risk and an area where AI can provide defensive value, evidenced by programs like CVP and its broader responsible scaling policy. Yet operationalizing nuanced trust distinctions at scale, especially under time pressure (the poster notes access to the new model expires July 7), remains technically and organizationally difficult. Classifiers must make probabilistic judgments about intent from text alone, and organizations often lag in building feedback loops that let verified partners bypass blunt-instrument restrictions.

More broadly, this incident highlights the tension AI companies face between speed of iteration and safety infrastructure maturity. Frontier labs frequently ship new models with conservative default guardrails, then iterate on classifier precision based on real-world feedback and complaints exactly like this one. The episode also underscores a recurring theme in the AI safety and security research community: legitimate defensive security work increasingly depends on frontier AI tools, meaning restrictive moderation can inadvertently hamper the very research meant to make AI systems and the broader internet safer. As agentic AI systems become more prevalent in cybersecurity tooling, expect continued friction—and continued pressure from labs like Anthropic to refine allowlisting, verification, and appeals processes for trusted research partners rather than relying solely on keyword-triggered automated downgrades.

Read original article →