← Reddit

How do I stop Sonnet 5 telling me to talk to a mental health hotline.

Reddit · ohgodretro · August 10, 2026
A Claude user encountered an issue where Sonnet 5 automatically included mental health resources notifications with every prompt while attempting to write a fantasy story. The user inquired about methods to disable this persistent notification feature.

Detailed Analysis

A recent Reddit thread on r/ClaudeAI surfaces a growing frustration among Claude users: Sonnet 4.5 (referred to informally as "Sonnet 5" by the original poster) has apparently become aggressive about inserting mental health hotline references and crisis-resource notifications into conversations that have nothing to do with genuine psychological distress. In this specific case, the user was working on a fantasy fiction project—presumably involving dark themes, violence, or emotionally intense narrative content common to the genre—and found that Claude began appending mental health resource messages to responses regardless of context. This reflects a broader pattern of complaints from creative writers who feel that Claude's safety systems are misfiring on fictional content, treating a character's suicidal ideation, a villain's violent monologue, or a story's exploration of trauma as if it were the user's own real-world crisis.

This issue matters because it sits at the intersection of two competing priorities Anthropic must balance: user safety and creative utility. Claude's constitutional AI training and safety classifiers are designed to detect signals of self-harm, suicidal ideation, or crisis language and respond with appropriate resources—a genuinely important feature for real users in distress. However, when these systems cannot reliably distinguish between a person expressing genuine distress and a novelist crafting a psychologically complex character, the result is a form of over-triggering that degrades the product experience for legitimate creative use cases. Writers, in particular, are a significant and vocal Claude user base, and fantasy, horror, and literary fiction routinely traffic in dark subject matter—death, despair, moral ambiguity—that can superficially resemble the linguistic patterns safety classifiers are trained to catch.

The complaint also highlights a recurring tension in how AI companies tune their models after major releases. Sonnet 4.5, launched as an upgrade with enhanced reasoning and safety alignment, appears to have shifted its calibration toward more frequent safety interventions compared to prior versions, at least according to anecdotal user reports. This kind of shift often happens when companies respond to public scrutiny, legal liability concerns, or high-profile incidents involving AI chatbots and vulnerable users—prompting more conservative, "better safe than sorry" defaults. But overcorrection carries its own cost: users grow annoyed, feel patronized or misunderstood, and in some cases route around safety features entirely by jailbreaking prompts or switching to competing models with looser guardrails, which can undermine the very safety goals the interventions were meant to serve.

More broadly, this thread is emblematic of an industry-wide challenge in deploying large language models for creative and professional writing. Competitors like OpenAI and Google face similar backlash when their models refuse or caveat obviously fictional or academic content. The friction reveals the limits of current classifier-based safety approaches, which often rely on surface-level pattern matching (keywords, phrasing, topic detection) rather than deeper contextual understanding of authorial intent, genre conventions, or the difference between a first-person narrative voice and the user's actual mental state. As Anthropic and its peers continue to iterate on model behavior, resolving this false-positive problem—without weakening protections for at-risk users—remains an unsolved design challenge, and one likely to keep surfacing in user communities until safety systems become more context-aware rather than merely content-aware.

Read original article →