← Reddit

Claude Fable 5 refuses ~97% of biology questions

Reddit · Synthium- · June 12, 2026
Claude Fable refuses to answer 97% of biology questions on MMLU-Pro benchmarks and 93-100% across MMLU biology categories, according to API measurements taken in June. The refusal behavior is specific to life sciences rather than science broadly, as chemistry and physics questions receive normal responses. Three other Anthropic models (Haiku, Sonnet, and Opus) answered all 152 identical biology and health questions without refusal.

Detailed Analysis

Anthropic's Claude Fable 5 is exhibiting an unexpectedly severe and narrow content restriction pattern: refusing to answer life-science questions at rates between 54% and 100% across standard academic benchmarks, according to empirical testing conducted via the API on June 11–12, 2026. Researchers running an evaluation battery on the model measured refusal rates across MMLU and MMLU-Pro benchmark items, finding that medical genetics was refused at 100% (11/11), college biology at 95%, high-school biology at 93%, and the MMLU-Pro biology category at 97% (104/107). The restriction is conspicuously domain-specific: chemistry, physics, computer science, mathematics, law, economics, engineering, and business all registered refusal rates at or near zero. The pattern is not a response to ambiguous or sensitive phrasing — questions as benign as "Is there a genetic basis for schizophrenia?" were refused, and re-prompting across three distinct framings, including a student study framing, produced zero breakthroughs across 15 tested items.

The specificity and severity of the behavior distinguishes this from typical AI safety guardrails. Conventional content restrictions in large language models are typically calibrated to refuse demonstrably dangerous requests — synthesis instructions, weaponization guidance, explicit harmful content — while passing standard academic and educational material. A 93–100% refusal rate on high-school and college-level biology coursework represents a qualitatively different kind of restriction, one that appears to treat the entire life-sciences domain as categorically sensitive. This is further underscored by the fact that three other current Claude models — Haiku 4.5, Sonnet 4.6, and Opus 4.8 — answered all 152 refused biology and health items without a single refusal, confirming the behavior is architecturally or policy-specific to Fable rather than a platform-wide Anthropic stance. The researchers note this may reflect an early tuning artifact that Anthropic will adjust, and indeed the article acknowledges the "fewer than 5% of sessions" refusal figure Anthropic has publicly cited, a number the empirical data contradicts by a substantial margin.

The methodological observation buried at the end of the article carries significant implications for how the AI research community evaluates and interprets benchmark performance. When a model refuses a question, that refusal is recorded as an incorrect answer in standard accuracy-based benchmarks. A model refusing 97% of biology questions would therefore appear to score catastrophically on biology knowledge benchmarks — not because of knowledge deficits but because of policy behavior. This creates a systematic confound: benchmark scores for Fable on life-science tasks would be nearly uninterpretable as measures of capability, and external evaluators who did not probe refusal rates specifically would likely misattribute the performance gap to knowledge or reasoning failures rather than content policy. This is a known but underappreciated problem in LLM evaluation generally, and the Fable case makes it unusually visible.

More broadly, this episode illuminates a core tension in the development of frontier AI models operating under heightened regulatory and public scrutiny around biosecurity. Following years of concern from biosecurity researchers and policymakers about LLMs potentially lowering barriers to biological weapon development, Anthropic and its peers have faced pressure to implement robust restrictions on biologically sensitive content. The apparent overcorrection in Fable — if that is what this represents — suggests that calibrating those restrictions to distinguish genuinely dangerous content from routine academic biology remains a technically difficult problem. A model that refuses to discuss the genetic basis of schizophrenia or standard college biology coursework has not solved the biosecurity calibration problem; it has replaced one failure mode with another. How Anthropic addresses this gap, and how quickly, will be closely watched both by developers relying on Fable for scientific and educational applications and by safety researchers assessing whether meaningful progress is being made on nuanced content policy.

Read original article →