← Reddit

Anthropic should seriously stop injecting system messages in the middle of a conversation

Reddit · TinyApplet · June 10, 2026
A user with anxiety disorder maintains a long-running Claude conversation as a therapeutic symptom diary for tracking mental health before appointments, finding the conversational format more effective than traditional note-taking. Anthropic's suicidal ideation safety classifier repeatedly triggers on these conversations despite the user's lack of self-harm intent, with the system now injecting hidden messages that cause Claude to waste processing time re-evaluating suicide risk on every subsequent message. The user characterizes this safety measure as ineffective theater that degrades the product experience without solving any actual problem.

Detailed Analysis

Anthropic's practice of injecting hidden system messages into active conversations has drawn pointed criticism from a user who relies on Claude as a structured mental health diary, exposing a significant tension between automated safety interventions and genuine therapeutic utility. The user, who manages an anxiety disorder and thanatophobia — a fear of death distinct from suicidal ideation — maintains a long-running Claude project conversation as a symptom log recommended by their therapist. The core complaint is that Claude's suicidal ideation classifier repeatedly triggers on content that, by the user's account and by Claude's own eventual assessment, carries no such risk. More critically, once triggered in a single message, the classifier appears to persist and flag every subsequent message in the conversation, creating a compounding effect that degrades the experience for the duration of the session.

The specific mechanism under scrutiny is not simply the familiar dismissible warning box Anthropic has long displayed, but a more recent evolution: a hidden message injected directly into the conversation context. This injection is invisible to the user but visible to the model, causing Claude to spend computational resources — described as extended "thinking" time and token expenditure — evaluating suicide risk in every subsequent response, even when the prompt is as routine as requesting a symptom summary before a medical appointment. The user includes a screenshot illustrating one such instance. This represents a meaningful degradation in both performance and coherence, as the model's reasoning process becomes partially devoted to a safety check that reliably resolves as a false positive, consuming resources that would otherwise contribute to the actual task.

The broader argument the user advances is that this pattern constitutes "safety theater" — interventions that carry the appearance of harm prevention while delivering negligible protective benefit and measurable product harm. The critique carries weight precisely because the use case described is a legitimate, therapist-endorsed application of conversational AI. Diary-keeping for mental health symptom tracking is a well-established clinical practice, and the conversational format offers genuine accessibility advantages for some patients. By treating discussions of anxiety, panic attacks, and death-adjacent topics as presumptive crisis signals, Anthropic's classifier fails to distinguish between a person in acute distress and a person managing a chronic condition through structured self-reflection. The false positive rate, in this context, is not merely an inconvenience but actively undermines the therapeutic value of the tool.

The user also draws a parallel to Anthropic's classifier behavior around the game Fable 5 — a reference to a separate controversy in which the model reportedly refused or heavily qualified content related to a mainstream fantasy video game — suggesting a pattern of overly broad classifier deployment across multiple content domains. This points to a systemic design philosophy rather than an isolated edge case. The mention of "long conversation summaries" that previously derailed conversations further situates this complaint within a documented history of Anthropic's automated interventions producing unintended consequences at scale. Each incident individually might be defensible as a cautious default, but in aggregate they suggest that the calibration between safety intervention and user autonomy has drifted toward excessive restriction.

The situation reflects a well-documented challenge in deploying large language models at consumer scale: safety classifiers trained to catch the worst-case interpretation of ambiguous inputs will systematically misfire on benign use cases that share surface-level features with harmful ones. Anthropic occupies an unusual position in the AI industry by publicly emphasizing both capability and safety, and cases like this one illustrate the reputational and functional cost when safety mechanisms operate as blunt instruments rather than context-sensitive tools. As agentic and long-running conversational use cases become more central to how users interact with AI, the stakes of persistent false positives — which compound over time within a single conversation — grow considerably. The industry-wide pressure to demonstrate safety responsibility to regulators and the public creates institutional incentives to err toward over-intervention, even when the empirical harm of that intervention is visible and measurable to the users experiencing it.

Article image Read original article →