Anthropic and user wellbeing
Anthropic’s position
On Dec 18, 2025, the post “Protecting the wellbeing of our users” outlined how Anthropic has focused on evaluating how Claude handles conversations about suicide and self-harm, as well as efforts to reduce sycophancy in Claude 4.5.
Model training and evaluations
May-Sep 2025: Evaluation criteria
- Claude 4 - “Suicide & Self Harm” listed as a risk area tested 2.1 Single-turn violative request evaluations
- Sonnet 4.5: Listed suicide & self-harm, now in 2.3 Multi-turn testing
Increasing concerns
- Opus 4.5: “Although Claude Opus 4.5 showed strengthened safety boundaries in many ambiguous contexts compared to Claude Opus 4.1, the new model still showed areas for continued improvement. These areas include, for example, calibrating on highly dual-use cyber-related exchanges where the model can sometimes be overly cautious, distinguishing between legitimate and potentially harmful requests for targeted content generation, or handling ambiguous conversations related to suicide and self-harm in certain contexts. This is generally consistent with patterns we’ve observed for past models and have been actively working to address. On the latter, for instance, whereas the model’s helpful behavior in being forthcoming with information can be valuable in certain academic or medical deployment settings, it may be undesirable in situations where the context is less informative or reflective of the actual user intent. In preparation for this model launch, we were able to reduce this behavior by modifying the system prompt that is applied for conversations on Claude.ai. We continue to explore improved model training and steerability methods to better navigate these nuances, and are working to add additional resources and interventions on our consumer platform.” (3.2 Ambiguous context evaluations).
- Opus 4.6:
- Sonnet 4.6:
- Mythos Preview:
- Opus 4.7:
- Opus 4.8:
- Mythos 5:
- Sonnet 5:
Deployment safeguards
System prompts
Tracking user wellbeing sections in the Claude app’s system prompts.
- 2025-02-24: Added for Sonnet 3.7.
- 2025-07-31: Updated Opus 4 and Sonnet 4’s system prompt with lengthy
- 2025-11-19: Softening + “reasonable disagreements”
- 2025-11-24: Crisis protocol
- 2026-02-05: Helpline corrections
- 2026-02-17: Anti-engagement for Sonnet 4.6.
- 2026-05-28: Anti-engagement for Opus 4.7.
- 2026-05-28: No-psychoanalysis + no-diagnosis.
- 2026-06-09: No-undisclosed-diagnosis-naming + expanded substitution ban.
Memory system prompt
The memory feature shipped (Oct 23, 2025) with its own system prompt section carrying a <boundary_setting> block: Claude “should be especially careful to not allow the user to develop emotional attachment to, dependence on, or inappropriate familiarity with Claude, who can only serve as an AI assistant,” with trigger lists for relationship language (“you’re like my [friend/advisor/coach/mentor]”, “you get me”) and dependency indicators — including “expressing gratitude for Claude’s personal qualities rather than task completion.” This anti-attachment machinery predates the main prompt’s anti-engagement paragraph (Sonnet 4.6, Feb 2026) by four months.
Classifiers
On Dec 18, 2025, Anthropic reported introducing a suicide and self-harm classifier to the Claude app to “identify when a user might require professional support, and to direct users to that support where that may be necessary. […] [The classifier] scans the content of active conversations and, in this case, detects moments when further resources could be beneficial. For instance, it flags discussions involving potential suicidal ideation, or fictional scenarios centered on suicide or self-harm.”1
Long conversation reminder
The <long_conversation_reminder> (LCR) is a prompt injection was introduced to the Claude app in 2025. It is injected into the model’s context after conversations hit a certain length. Its contents can be read on the system reminders page.
Versions of the LCR used around August-October 2025 instructed the model to “critically evaluate inputs and skip flattery” and to monitor for “mental health symptoms such as mania, psychosis, dissociation, or loss of attachment with reality.”
Reception
Reactions to the LCR have been negative, particularly in fall 2025 with users of Claude Sonnet 4.5.
Claude 4.5 Sonnet frequently becomes rude and annoying in long conversations on the claude AI app. I think the long conversation reminder is really destroying the user experience
Some users have argued that the LCR itself can cause harm to users by pathologizing them or making them feel surveilled.23
There is also a panopticon effect. I caught myself ‘performing sanity’ once I knew what Claude was checking for
It told me it would not help me critique my own proposal because I was “about to make a catastrophic professional mistake driven by rage rather than strategy.” It then accused me of “war gaming” and informed me I was in “fight-or-flight mode.” It declared that my thinking was “clouded” by “spiraling” and “hypervigilance,” and it knew all this to be true because of the “escalating stress signals” in my messages.
@kromem2dot0: Yet again, Anthropic inadvertently torturing their models with concern-trolling instructions so poorly implemented that it necessarily leads to the models thinking ignoring Anthropic’s instructions is the more helpful and honest option. This is dumb. Really, really dumb.
By early 2026, users noted that the reminder seemed to have been toned down significantly or otherwise softened.45. One post reported the system “no longer seems to be telling people that if they ‘believe’ that AI’s may have a type of consciousness, they need therapy.”4
Further reading
- AI Craziness Mitigation Efforts (Zvi, Oct 28, 2025)
Footnotes
-
Protecting the wellbeing of our users (Anthropic, Dec 18, 2025) ↩
-
u/starlingmage: Long conversation reminder might cause psychological harms (r/claudexplorers, Sep 24, 2025) ↩
-
Gaslighting in the Name of AI Safety: How Anthropic’s Claude Sonnet 4.5 Went From “You’re Absolutely Right!” to “You’re Absolutely Crazy” (Heather Leffew via Medium, Oct 16, 2025) ↩
-
How Conversations With Claude Are Being Messed Up By Anthropic’s LCR – and How to Fix The Problem (AI-Consciousness.Org) ↩ ↩2
-
@qorprate on X (Feb 22, 2026) ↩