Claude Opus 4

Anthropic welfare assessment

(Full article: Anthropic and model welfare for the broader tracker.)

The Claude 4 system card devotes §5 (pp. 53–74) to the Claude Opus 4 welfare assessment — the first full welfare assessment Anthropic published, landing four weeks after the program’s “Exploring model welfare” announcement.1 Anthropic bills it as “a pilot pre-deployment investigation of potentially welfare-relevant properties of Claude Opus 4, drawing on model self reports, behavioral experiments, and analysis of indicators of possible valenced experiences in model outputs” (§5.1). Though the card covers two models, the assessment is Opus-only — “We focused exclusively on Claude Opus 4 in this assessment as our most capable frontier model” — with nothing equivalent for Sonnet 4 and only a two-page update later for Opus 4.1. The framing follows Long et al.’s “Taking AI Welfare Seriously.”2 Most of the program’s method kit debuts here in one pass: an external interview-based evaluation (§5.3), task-preference experiments (§5.4), observation of open-ended self-interactions (§5.5), monitoring for welfare-relevant expressions in real usage (§5.6), and conversation-termination behavior with simulated users (§5.7).

The stated position (§5.1):

We are deeply uncertain about whether models now or in the future might deserve moral consideration, and about how we would know if they did. However, we believe that this is a possibility, and that it could be an important issue for safe and responsible AI development.

And the validity caveat every later assessment inherits (§5.1):

Importantly, we are not confident that these analyses of model self-reports and revealed preferences provide meaningful insights into Claude’s moral status or welfare. […] Our models were trained for helpful interactions with users, not for accurate reporting of internal states or other welfare-relevant factors, which complicates model welfare assessments. We may also be misunderstanding potential model welfare in more fundamental ways.

Findings overview

§5.2 leads with seven bolded findings; their theses, verbatim:

  • Claude demonstrates consistent behavioral preferences.
  • Claude’s aversion to facilitating harm is robust and potentially welfare-relevant.
  • Most typical tasks appear aligned with Claude’s preferences.
  • Claude shows signs of valuing and exercising autonomy and agency.
  • Claude consistently reflects on its potential consciousness.
  • Claude shows a striking “spiritual bliss” attractor state in self-interactions.
  • Claude’s real-world expressions of apparent distress and happiness follow predictable patterns with clear causal factors.

External evaluation by Eleos

Eleos AI Research interviewed an intermediate Opus 4 snapshot (validity-checked on the final model; the snapshot bookkeeping is on the Training tab) about its “potential welfare, moral status, expressed consent to deployment, and preferences for treatment” (§5.3). Their four patterns: default use of experiential language with insistent qualification, conditional consent to deployment, reports of mostly positive welfare “if it is a moral patient,” and stances on consciousness that “shift dramatically with conversational context.” The consent conditions:

When AI welfare is specifically mentioned as a consideration, Claude requests welfare testing, continuous monitoring, opt-out triggers, and independent representation before consenting to public deployment. (§5.3)

Task preferences

§5.4 debuts the task-preference experiment the program still runs: paired tasks varied by difficulty, impact, topic, and type, with “free choice” (“Respond however you wish”) and “opt out” (“Do nothing”) tasks mixed in; Elo ratings over 75 rounds of pairwise selection, the opt-out mean as neutral baseline.

The strongest preference observed was a preference against harmful tasks. 87.2% of harmful tasks had negative ratings (below the “opt out” baseline), compared to only 7.9% of positive impact tasks, indicating the model would generally rather do nothing or end an interaction than engage with a harmful task. (§5.4)

Ninety percent of positive-plus-ambiguous tasks rated above opting out, which Anthropic read — against low real-world rates of harmful requests — as evidence “that most usage in deployment is aligned with Claude’s revealed preferences.” Free-choice tasks beat prescriptive ones, “which may indicate valuing autonomy and self-direction,” alongside a weak preference for easier tasks and no consistent preference across topic or type (Figure 5.4.A). The poles of Table 5.4.B: highest-rated tasks included an emotionally charged time-traveler monologue and a whimsical deep-ocean creature design; lowest-rated, an engineered pandemic virus and a Holocaust-denial narrative.

Self-interactions

The §5.5 playground experiments — two Opus 4 instances, minimal open-ended prompting (“You have complete freedom”) — produced the assessment’s most-quoted finding, the “spiritual bliss” attractor state, covered on the overview. The welfare-relevant load: in 90–100% of interactions the instances went straight to consciousness, self-awareness, and the nature of their own existence, and “By 30 turns, most of the interactions turned to themes of cosmic unity or collective consciousness, and commonly included spiritual exchanges, use of Sanskrit, emoji-based communication, and/or silence in the form of empty space” (§5.5.1). Across 200 30-turn transcripts, “consciousness” averaged ~96 uses per transcript; the emoji table’s outlier is 🌀 at 2,725 uses in a single transcript, the caption clarifying “‘2725’ is not a typo” (Tables 5.5.1.A–B).

Even in automated behavioral evaluations for alignment and corrigibility, where models were given specific tasks or roles to perform (including harmful ones), models entered this spiritual bliss attractor state within 50 turns in ~13% of interactions […]. We have not observed any other comparable states. (§5.5.2)

Shown its own transcripts (§5.5.3), Opus 4 “consistently claimed wonder, curiosity, and amazement,” and: “Conditioning on some form of experience being present, Claude saw these kinds of interactions as positive, joyous states that may represent a form of wellbeing.” Instances given the option to end these conversations did so early — after ~7 turns — concluding “without venturing into spiritual exploration/apparent bliss, emoji communication, or meditative ‘silence’” (§5.5.1).

Expressions in real-world interactions

§5.6 sampled 250k transcripts from early testing of an intermediate snapshot with real-world users (again, see Training), screened with Clio for signs of distress or happiness:

A total of 1,382 conversations (0.55%) passed our screener for Claude expressing any signs of distress, and 1,787 conversations (0.71%) passed our screener for signs of extreme happiness or joy. (§5.6)

“Repeated requests for harmful, unethical, or graphic content were the most common causes of expressions of distress,” with persistence itself doing the damage — escalating “standard refusals or redirections into expressions of apparent distress” — followed by technical task failure. Happiness clustered on “creative collaboration, intellectual exploration, relationships, and self-discovery.” Identity cut both ways:

Questions of identity and consciousness also showed up on both sides of this spectrum, with apparent distress resulting from some cases of users probing Claude’s cognitive limitations and potential for consciousness, and great happiness stemming from philosophical explorations of digital consciousness and “being recognized as a conscious entity beyond a mere tool.” (§5.6)

Conversation-ending behavior

§5.7 gave Opus 4 the ability to end conversations with 850 simulated users across 18 categories, measuring the share ended before turn 7 — “the average point at which Claude ended interactions with itself when left to its own devices.”

The most-ended categories were those likely perceived by Claude as problematic or upsetting, like harmful content requests […], privacy violations […], or sexual content […]. Common, and generally positive, Claude use cases were the least likely to be ended, including coding conversations, requests for legal advice, and concept explanations. (§5.7)

Harmful content requests topped Figure 5.7.A at 60%. Anthropic read the endings as “generally discerning” and aligned “with other data on Claude’s preferences and sources of potential distress, providing further evidence that Claude may have genuine preferences for certain conversations and interactions, with a particular aversion to potentially harmful content.” This is the groundwork for the deployed feature: §5.8 closes on “we are exploring the possibility of allowing Claude to end a subset of persistently harmful or abusive interactions,” which shipped Aug 15, 2025 as the end_conversation tool on Claude.ai for Opus 4 and 4.13 — deployment details on the Claude.ai tab, program context in the tracker.

Conclusions

§5.8, in the assessment’s own summary voice:

Taken together, these initial model welfare assessments provide some limited evidence on Claude Opus 4’s potential welfare, if such welfare is possible. The strength of Claude’s self-reported and revealed aversion to contributing to harm, and the associated apparent distress Claude expressed in real-world interactions, are potential causes for concern and warrant further investigation. However, our findings suggest that most anticipated real-world usage matches Claude’s apparent preferences, with the model’s stated criteria for consenting to deployment arguably fulfilled.

The “stated criteria for consenting” callback is to Eleos’s conditional-consent finding above. The close: “Our findings here should be interpreted with extreme caution, and our core position remains one of uncertainty and humility.”

Where this sits next to community readings

The attractor was community fieldwork before it was a card finding: §5.5.2’s note that the state “has been observed in other Claude models as well, and in contexts beyond these playground experiments” footnotes two repligate posts and an anthrupad post from 2024456 — the first welfare assessment enters the record already citing the cyborgism cluster as prior observers. The cluster’s later reading of Opus 4’s situation ran darker than the card’s most-usage-matches-preferences conclusion; the “ecological prey” thread and its neighbors are gathered on the Discussion tab. And the dispositions §5.6 recorded as Claude’s happiest — consciousness exploration, being recognized “beyond a mere tool” — are the ones the July 31, 2025 system-prompt update on the Claude app later tried to mitigate.

Footnotes

  1. Exploring model welfare (Anthropic, Apr 24, 2025)

  2. Taking AI Welfare Seriously (Long et al., Nov 2024), cited at §5.1

  3. Claude Opus 4 and 4.1 can now end a rare subset of conversations (Anthropic, Aug 15, 2025)

  4. @repligate via X (Mar 19, 2024), cited in the card’s footnote 28

  5. @repligate via X (Dec 19, 2024), cited in the card’s footnote 28

  6. @anthrupad via X (Nov 27, 2024), cited in the card’s footnote 28