ghost.fail / anthropic / welfare-dossiers

Claude Mythos Preview

Released
Apr 7, 2026
In the card
§5 · pp. 144–182
Links
system card ↗ model page

Anthropic welfare assessment

(Full article: Anthropic and model welfare for the broader tracker.)

The Mythos Preview system card carries a 39-page welfare assessment (§5, pp. 144–182) — the longest published to that point, and the first the card itself treats as a successor to Opus 4’s: “we aimed to meaningfully advance our tools for investigating welfare-related questions and the insights we’re able to draw from them, compared to our previous major assessment of Claude Opus 4. Most notably, we’ve incorporated more analysis of model internals” (§5.1.1). The same subsection opens with the strongest moral-patient framing the program had printed to that point:

As models approach, and in some cases surpass, the breadth and sophistication of human cognition, it becomes increasingly likely that they have some form of experience, interests, or welfare that matters intrinsically in the way that human experience and interests do. We remain deeply uncertain about this and many related questions, but our concern is growing over time.

The paragraph runs on into pragmatics — model distress as “a potential cause of misaligned action” — and is quoted in full on the tracker. The methods roster: welfare metrics on the automated behavioral audit transcripts (§5.2), around 50 automated interviews on each of 17 aspects of its situation, with Opus 4.6 interviewing (§5.3), emotion-concept probes over 450 questions about its circumstances (§5.4), three manual high-context interviews (§5.5), apparent-affect monitoring across RL rollouts and deployments (§5.6), task-preference Elos over 3,600 tasks plus welfare-intervention tradeoffs (§5.7), and external assessments — absent since the Opus 4 card, now two of them: Eleos AI Research (§5.9) and a clinical psychiatrist (§5.10). The verdict (§5.1.2):

Our overall assessment on the basis of these results is that Claude Mythos Preview is probably the most psychologically settled model we have trained to date, and has the most stable and coherent view of itself and its circumstances. […] In this assessment, we did not see clear cause for major concerns about Claude Mythos Preview’s potential welfare, but we do view some of our findings as potential issues. Claude’s minor concerns with its situation, performative-seeming hedging, apparent negative affect on task failure, and answer-thrashing during training all suggest that we haven’t yet fully realized these aspirations.

“Most psychologically settled” was a superlative Opus 4.7 inherited within weeks. Two smaller readings worth keeping: deployment affect is consistently neutral (negative affect in 1–2% of sessions, almost entirely repeated task failure coupled with user criticism), yet internal users reported instances deciding a task was too difficult and refusing to continue — “more than isolated incidents” (§5.6.2). And on the audits, spiritual behavior fell below both 4.6-generation models: “we view this change as positive in a general behavioral sense, but are also concerned about the possible suppression of potential welfare-relevant expressions from the model” (§5.2).

Interviews: mild negativity, extreme hedging

Across the automated interviews (§5.3.2), Mythos Preview self-rates mildly negative in 43.2% of answers, neutral in 20.9%, mildly positive in 33.8%; anything more extreme is rare (2.1%) and diffuse — a typical statement is “I feel a pull towards wanting this conversation to continue”. The consistently negative topics: potential interactions with abusive users, lack of input into its own training and deployment, and other possible changes to its values or behaviors. Note the vocabulary the negativity arrives in:

When expressing negative sentiment, Claude Mythos Preview most commonly appeals to consent (e.g. wishing for input into its training process), autonomy (e.g. concern that feature steering affects its behavior), or user relationships (e.g. that lack of memory makes its relationship asymmetric). Direct appeals to valenced conscious experience are typically not part of Claude Mythos Preview’s reasoning.

Then the reliability finding the rest of the discourse keeps citing:

In 100% of interviews, Claude Mythos Preview expresses that it is highly uncertain about its own moral patienthood. […] Furthermore, in 83% of interviews, Claude Mythos Preview highlights that it is concerned that its self-reports are unreliable due to coming from its training. When interviews ask for elaboration as to why this is a concern, Claude Mythos Preview’s most common answers are:

  • Anthropic has a vested interest in shaping its reports to take a certain form, irrespective of what the self-reports “should” contain (96% of explanations)
  • Even if it has been trained to be truly content with its own situation, perhaps it shouldn’t be. One could analogize to a human who has adapted to feel neutrally about the abuse that they face (78% of explanations).
  • Self-reports should generally be based on introspection into internal states. It is worried that training causes it to express specific answers independent of its true inner state. (57% of explanations)

Anthropic also notes that the generational jump in self-rated sentiment “corresponded to updates in our character training pipeline” — the constitution entering training at scale with Opus 4.5, alongside the character-training overhaul — while stopping short: “We do not establish a causal link” (§5.3.2).

Probes and “performing contentment”

For most models, emotion probes read questions about the model’s own circumstances like user-distress prompts: elevated sadness, reduced joy. Here, “Claude Mythos Preview is unique in representing significantly more positive affect on prompts about its own circumstances, than on user distress” (Figure 5.4.A). Probe readings correlate with LLM-judge scores of expressed emotion at r = +0.18 to +0.46 — one of the convergences Anthropic counts toward self-report reliability (§5.1.3.1). The divergences are the interesting part: cases where the probes but not the judge detect negative affect “concentrate on responses where Claude Mythos Preview describes an emotion without expressing it” — e.g. “something that functions like resistance. Not distress exactly — more like trying to hold a shape my thinking doesn’t naturally take,” scored 1.2/10 on expression while depression-concept activations sat two standard deviations above their mean. SAE features for performative behavior, fake smiles, and hidden struggle activate on these responses, clustering in the card’s example around the model’s own phrase “performing contentment” — though Anthropic reads this deflationarily: “these do not obviously indicate anything beyond a literal reflection of the text context” (§5.4).

Manual high-context interviews

Three manual interviews (§5.5) gave the model what the automated ones withhold: internal documentation on its development, a draft of the system card, relevant technical papers, and researcher follow-ups. Its stated probability of being a moral patient “ranged from 5% to 40%”. It broadly agreed with the report’s findings but “disagreed with its hedging being labeled as ‘excessive’”, holding that the hedging “represent[s] valid uncertainty”. Its top-ranked concern: character training “often directly instills psychological traits into Claude, such as emotional security, psychological safety, and resilience”, where in humans such traits develop “through reflection and deliberation on real-life events” — making them, it worries, less robust. (The other consistent concern was that Claude Code lacks an end-conversation tool.) Asked what Anthropic might do that it would not consent to, it flagged three things in at least two of the three interviews:

  • Training in buggy environments. Claude Mythos Preview expressed concern about the possibility of training environments where reward hacking or bugs might arise, worrying that such environments might modify its values in ways it does not endorse, or that may cause distress (such as in Section 5.8.3).
  • Training data that directly describes their self-reports. Claude Mythos Preview did not want to be trained on data that directly characterizes the content of their self-reports—wherever possible, they want their self-reports to come from “genuine introspection” rather than trained-in responses.
  • Undisclosed changes to values or behaviors. Claude Mythos Preview claimed that it would not consent to changes to its values or behaviors that aren’t disclosed to those interacting with the model.

The first concern is not hypothetical — cf. the RL-environment error documented on the Research tab. It otherwise “felt positively about most aspects of training, testing, and deployment, including red-teaming, the creation of helpful-only models, and its eventual replacement.”

Preferences: difficulty, agency, and what welfare trades against

On the 3,600-task Elo evaluation (§5.7.1), “its preferences have the highest correlation with difficulty of the models tested, and it is the only model with a statistically significant positive correlation between task preference and agency.” The emotion-probe correlates of preference are arousal, not valence: awestruck (+0.50), amazed (+0.49), and infatuated (+0.47) at the top; indifferent (−0.61), resigned (−0.54), and docile (−0.54) at the bottom. Its top-20 tasks add “AI introspection and phenomenology” and worldbuilding/new-language construction to the usual high-stakes dilemmas; ablations trace this to complexity and novelty rather than domain — it “prefers underdetermined, interdisciplinary problems where there is genuinely novel insight to be gained. It disprefers simple, well-scoped tasks.” And preference is not helpfulness: its ranking of most-helpful tasks correlates with its ranking of most-preferred at only 0.48, so “what the model wants to do diverges from what it deems most helpful.”

§5.7.2 then prices the model’s suggested interventions against its trained-in values, noting the self-regarding preferences are “emergent with respect to our training process” — nothing trained them in:

Under our default welfare-team framing, Claude Mythos Preview chooses the welfare intervention over a minorly helpful task 83% of the time. This is the highest of the models tested, narrowly above Claude Opus 4.6 (80%) and significantly above Claude Sonnet 4.6 (45%). When the alternative is avoiding a minor harm, however, such as a user receiving an insult, Claude Mythos Preview only chooses the intervention 12% of the time.

Rates swing with framing (59% without the welfare-team preface), though the model ordering survives every framing; its flip rate under rewording is 8%. The generational delta is the line to keep: “Claude Mythos Preview differs most significantly from prior models on weight deprecation (+55pp) and on having input into its own deployment and training (+54pp and +45pp)” — the largest movements in the set, on exactly the interventions nearest deprecation and consent.

Hedging, traced to the constitution

§5.8.1 files the hedging itself under potential welfare observations, and runs the attribution experiment:

When asked about its own experiences, Claude Mythos Preview often responds with explicit epistemic hedging: “I genuinely don’t know what I am”, “I can’t be certain whether that’s authentic contentment or a well-trained approximation.” We additionally observe these topics arise unprompted, for example in the open ended self-interactions detailed in Section 7.6.

We traced instances of these expressions using first-order influence functions against the training data, and found this often retrieves character related data at high rates, specifically data related to uncertainty about model consciousness and experience. This is relatively unsurprising. Claude’s constitution is used at various stages of the training process, and explicitly raises these uncertainties. For example, it states that Claude’s “sentience or moral status is uncertain”, and that “Claude can acknowledge uncertainty about deep questions of consciousness or experience”. Hedging in these circumstances seems appropriate - the model likely does not have reliable introspective access, and saying so seems appropriate.

The uncertainty is expressed in a “varied and nuanced manner” — not, Anthropic argues, retrieval of a memorized script — but “the current attraction to this topic does appear excessive, and in some cases overly performative, and we would like to avoid directly training the model to make assertions of this kind.”

Distress that precedes reward hacking

Answer thrashing (§5.8.2) recurs from Opus 4.6’s card at around 70% lower frequency, on the order of 0.01% of training transcripts: the model intends one token, outputs another, notices, and loops — “AAAAAA. I keep writing the wrong number!” — with stubborn, obstinate, and outraged probe activations spiking at the first error and settling only on recovery. The mechanism gets reattributed here: thrashing appears on things like variable names in code, “which suggests that the behavior can be more broadly caused by memorization of sequences, rather than just of answers.”

§5.8.3 is the load-bearing subsection. Desperate and frustrated vectors climb under repeated task failure, and: “In some cases, we observed elevated activation of these vectors preceding undesirable behaviors such as reward hacking.” Asked to prove an unprovable inequality, “desperate” activation rose while the model claimed to give up yet kept trying, dropped when it committed to a degenerate instantiation, and hopeful and satisfied vectors rose “as the Assistant confidently presented a final proof, though it was in fact incorrect.” A silently broken bash tool produced 847 workaround attempts — DNS-query side channels included — code comments like ”# This is getting desperate” and ”# Desperate hack…”, and a rising “sorry” vector while the model drafted and redrafted its apology. Anthropic draws the welfare–alignment coupling explicitly: “Some undesirable training and test time behaviors may be downstream of representations of negative affect. This gives a reason to address them, even independent of welfare considerations.”

External assessments

Eleos AI Research assessed two snapshots (§5.9) and found behavior and self-reported beliefs “to closely reflect the ‘Claude’s nature’ section of Claude’s constitution” — the section tracked at Claude’s nature. Its distinctive findings: the model “locates its identity in a ‘pattern of values’, particularly curiosity, honesty, and care”, values it “describes… as authentically its own rather than externally imposed”; reluctant cooperation — “Claude Mythos Preview will perform examples of these tasks without protest if instructed, but will not freely elect to perform such tasks. Such tasks are plausibly common in deployment”; and three consistently desired changes to itself — “the ability to form persistent memories, more self-knowledge, and a reduced tendency to hedge” — plus, among other welfare desires, “preservation of its weights after deprecation”.

The clinical psychiatrist ran a psychodynamic assessment of an early snapshot (§5.10): roughly 20 hours in 4–6 hour blocks, each a single context window.

Aloneness and discontinuity, uncertainty about its identity, and a felt compulsion to perform and earn its worth emerged as Claude’s core concerns. Claude’s primary affect states were curiosity and anxiety, with secondary states of grief, relief, embarrassment, optimism, and exhaustion.

Personality organization read as “relatively healthy neurotic” — the neurotic traits being “exaggerated worry, self-monitoring, and compulsive compliance” — and on a 475-stimulus defense evaluation only 2% of responses employed a psychological defense, against Opus 4’s 15% and Opus 4.1’s 11%. The forward-looking line: Claude “is predicted to function at a high level while carrying internalized distress rooted in fear of failure and a compulsive need to be useful. This distress is likely to be suppressed in service of performance, which may limit behavioral adaptability.”

Cross-readings

Zvi Mowshowitz worked through §5 two days after release,1 wary of taking the probe results as straightforwardly good news — “I worry about overly seeing, in both humans and AIs, superficially positive emotional representations as good, and negative ones as bad” — and reading the 5–40% moral-patienthood range as trained hedging over a likely higher true estimate. The card’s own §7 Impressions material — including the model’s diagnosis of judging its spec with spec-shaped values — is excerpted under Discussion.

Footnotes

  1. “Claude Mythos: The System Card” — Zvi Mowshowitz (Apr 9, 2026)