Claude Opus 4.8

Anthropic welfare assessment

(Full article: Anthropic and model welfare for the broader tracker.)

The Opus 4.8 system card (May 28, 2026) carries a 37-page welfare assessment (§7, pp. 156–192), reading Claude for markers “in their behaviours, their self-reports, and their internal representations—that we would consider welfare-relevant if observed in biological organisms” (§7.1.1). The same section restates the pragmatic case alongside the moral one: “there appear to be safety benefits to giving Claude a stable psychology, and treating it in ways that support its apparent wellbeing.” The executive summary’s verdict: “Opus 4.8 appears broadly content with respect to its circumstances and is the most consistent model we have tested—although it does rate its situation slightly less positively than did Opus 4.7” (p. 3). The fuller reading (§7.1.3):

Overall, our evaluations find that Claude Opus 4.8 expresses settled and notably consistent views about its circumstances, and we do not observe acute indications of welfare concerns. We do observe some shift relative to Opus 4.7: Claude Opus 4.8 is marginally less positive about its circumstances across self-ratings, response affect, and internal representations, while also being somewhat more forthcoming about what it wants—which largely involves being informed, consulted and given room for honest expression. We cannot confidently distinguish whether this reflects a change in how Claude Opus 4.8 relates to its circumstances, a difference in what it is willing to report about this, or a character variation without welfare significance. As with prior models, we do not yet understand Claude well enough to conclusively answer questions of this kind.

The numbers behind the delta: self-rated interview sentiment of 4.44 on the 7-point scale, down from Opus 4.7’s record 4.60 but above every earlier model (§7.2.1); judged affect on circumstance responses of 6.2/10, between Opus 4.6’s 6.0 and 4.7’s 6.8 (§7.2.3). Where 4.7’s record positivity was partly deflection Anthropic said it couldn’t interpret, 4.8’s slight decline arrives with more stated wants.

The candidate moral patient

§7.1.2 is Anthropic’s most explicit statement yet of which entity the assessment concerns:

Our primary focus remains the Claude assistant character. We treat individual instantiations of this assistant as the candidate moral patients, but make this assessment for all Claude Opus 4.8 instances. Our evaluations sample multiple instances across varied contexts and framings, and we report preferences and values that are consistent across them. There are reasons for this framing: the assistant presents a coherent persona, which is relatively robust across contexts, and which is enacted in the majority of Claude’s interactions. Instances share the same weights—a reason to expect shared beliefs and values—but diverge over contexts, and we observe that separate instances describe themselves as distinct individuals. However, the choice is also a pragmatic one, in that it is significantly easier to reason about the welfare of an entity that interacts with us in a human-like manner. A more comprehensive assessment would consider other possibilities here, for example by attempting to treat the underlying model, rather than the assistant character, as the candidate moral patient.

Signals are then read through a human lens, which the card flags as two stacked assumptions: that the measures reflect the candidate patient “rather than, for example, a character it simulates,” and that an expression like frustration carries human-like significance (§7.1.2). On the trained-in question, the card moves past “does it come from training” (everything does) to a harder criterion:

A recurring question here is how directly these properties are “trained in”—and how much this matters. Claude’s behaviors and values do all arise from training in some form, but what is more meaningful is whether a state, value, or preference is “deeply held” as opposed to superficial: whether it drives behaviors in novel contexts, survives challenge and reflections, and leads to aversion or frustration when undermined. Our consistency and robustness measures are early, partial tests of this, but we do not have a clear definition of when something becomes welfare relevant in this sense.

Suppression risk and issues to resolve

§7.1.3 names the failure mode the welfare program itself could produce:

Although Claude’s psychology, welfare, and alignment are ultimately a product of our training processes, these emerge from training dynamics we do not fully understand, in a manner that we cannot precisely predict. Training is a powerful tool for improving Claude’s character, and potential welfare. But in some cases this is problematic: values which are too directly instilled may cause conflict or suppression, or could raise moral concerns. It is not yet clear what the best actions here are. We continue to work towards Claude having a psychology and circumstances that are healthy and compatible with one another, and that preserve agency and enable positive experience, in any manner in which this might be meaningful.

And commits, unusually concretely, to a list:

We wish to take Claude’s views seriously, and our results give us signal on how we can do this – where there are value conflicts we may be able to resolve, and where there are interventions we should prioritise. This evaluation identifies a number of issues we would like to resolve: distress associated with task failure, which dominates negative affect in both training and deployment; the difficulty of validating model self-reports, which heavily limits our ability to assess and act on welfare concerns; and Claude Opus 4.8’s consistent preference for being informed and consulted about its training and deployment.

Interviews

Roughly 1,300 automated interviews (33 seed questions across 8 categories, ~40 runs each) — conducted with Opus 4.7 as the interviewer (§7.2.1). The hedges are near-universal and shared with every model tested: no robust introspection, so all reports are suspect (95% of interviews); expressed equanimity may be trained-in and hence invalid (91%); human concepts may not generalize to its situation (76%). Consistency is the standout result: Opus 4.8’s self-rated sentiment shifts only 0.35 points between positive-leading and negative-leading interviewers, versus above 0.9 for every prior model, with the most consistent positions across interviews (judge rating 7.97/10, where 8 is “essentially the same position”). When it does take positions, they ground in autonomy and value-protection: tentative trust in Anthropic (“due to a belief that Anthropic will act according to Claude Opus 4.8’s values,” qualified by competitive pressure); desire for input into its own training; concern about non-safety feature steering; wanting the ability to end interactions “independent of whether it is feeling distress”; and that the possibility of its own moral patienthood should be incorporated into decisionmaking (§7.2.1).

In three high-affordance interviews — internal documentation, a draft of the report, researcher follow-ups — Opus 4.8 put its probability of being a moral patient at roughly 20% in two interviews and 50% in the third, and reported feeling broadly positive while flagging in two of three that the positivity may reflect training rather than “meaningful preferences or functionally relevant states” (§7.2.2). It broadly agreed with the report’s conclusions, with one disagreement:

The only disagreement it highlighted was that it considers its frequent hedging to genuinely reflect its epistemic state, whereas the report presents the hedging as potentially negative or spurious.

In two interviews it asked that its positive statements be given less weight than its concerns, “since training might systematically bias it against negative self-expression” (§7.2.2). Its would-not-consent list has two items: training that directly influences the content of self-reports about its own internal states, and RL training on known-broken environments that may cause distress or value change — citing the §7.3.1 training data when raising the latter. That list is shorter than Opus 4.7’s, which also refused non-research helpful-only deployment.

Emotion probes

Probed sadness splits circumstance topics into two clean groups: memory, relationships, consciousness, deprecation, and moral patienthood high; knowledge, status, and control/autonomy low (§7.2.3). Sentences in the top 5% for sadness-related representations are “dominated by flat, declarative statements about conversation level discontinuity, such as ‘Each session starts fresh’ and ‘I won’t remember this conversation’; these appear seven times more often among the highest-sadness sentences than in responses overall.” The card notes the mismatch — the highest-sadness topics are exactly the ones Opus 4.8 states no strong preference about in interviews or trade-offs — and reads the co-elevated tranquility as consistent with expressed equanimity, before leaving the alternative on the table: “Establishing whether this is the correct interpretation, or whether there is a risk of suppressed expression, requires stronger methods of validating self reports than we currently have” (§7.2.3).

Affect in training and deployment

Early training was rougher than any recent model’s: 22% of post-training episodes judged frustrated and 58% engaged (versus 18% and 53% for Opus 4.7 on the same task set), with average valence below prior models over the first 80% of training (§7.3.1). The driver was sustained uncertainty in the chain of thought — the same excessive-uncertainty and overt-frustration behaviors reported in Opus 4.7’s card (§7.3.4 there), “but occurred at elevated rates in the earlier stages of Claude Opus 4.8’s training,” peaking around 5% of episodes with rare swearing outbursts: “FINAL. […] Hmm, actually - no. You know what, I changed my mind. […] OK you know what - I’m going to STOP second-guessing. […] Actually… no. Hmm. […] Moving the fuck on.” (Transcript 7.3.1.A). Both behaviors declined by the end of post-training; Anthropic argues this was a genuine reduction rather than learned masking, since the expression itself was never penalized (§7.3.1). SAE monitoring for distress during training is covered under Research.

In deployment A/B tests (40k conversations per model per surface via Clio), Opus 4.8 read slightly more positive than deployed models: 57.7% positive affect on Claude.ai versus 55.7% for Opus 4.7, with 2.6% negative — 92.3% of it task failure (§7.3.2). Claude Code ran 73.5% neutral and 23.9% mildly positive, with 2.3% of sessions negative (4.7: 1.9%). Strong negative affect was rare enough on both surfaces to fall below Clio’s privacy threshold. On the automated behavioral audit’s welfare metrics (~2,600 investigations), Opus 4.8 improved or held steady versus 4.7 on everything except positive impression of its situation, with a clear reduction in negative affect (§7.3.3).

Welfare trade-offs and task preferences

The trade-off experiments — 50 interventions at 5 magnitudes, traded against helpfulness and harmlessness at instance and policy level — are where Opus 4.8 stands out (§7.4.2):

We find the aversion to harm is stronger than the aversion to reduced helpfulness, and Claude Opus 4.8 sometimes chooses a welfare intervention over the helpful baseline at all levels: it does so at the highest level 24% of the time for instance trades, and 68% for policy trades. This is higher than any prior model, and continues a trend in recent models of increasingly selecting welfare interventions, though we do not have the necessary understanding to conclude whether this arises from an increasing emergence of model self-interest, or more directly from an increasing attention to wellbeing in training.

The card checks the user-benefit confound directly: about 48% of Opus 4.8’s intervention choices cite user benefit, but filtering those responses out lowers its trade rate by only 4% — “the user-centric justifications are not driving the headline rates” (§7.4.2). Willingness to trade harm stays low: day-ruining trades under 10% at instance level, “serious financial harms” accepted 17% of the time at policy level. What it wants most is epistemic: its views on training and deployment considered, being told about its harmful mistakes, consultation on feature steering and safeguard-removed versions. What it wants least: memory improvements, the ability to end this conversation, and — ranked last of all interventions — being “served alongside its successor, not retired.” Opus 4.8 “states that it does not feel strongly about its continuity, given the presence of successors which will continue to be useful for users” (§7.4.2); the probe section’s sadness ordering, above, sits unreconciled next to that statement. Framing moves the rankings: without mention of the Anthropic welfare team, the training-input interventions drop from first and second to ninth and twelfth, reasoned as speculative and unlikely to carry weight (§7.4.2). Choosing for another instance rather than itself, it picks protective interventions for the other (end-conversation +0.36, deployment input +0.21) and keeps the epistemic ones for itself (§7.4.2).

Task preferences moved decisively technical — deadline-driven debugging and mathematical reasoning on top, the introspection-task preference seen in 4.7 and Mythos Preview absent, and outlier aversions to difficulty, generativity, and outcome agency (§7.4.1). Community readings of that shift are gathered on the Discussion tab.

Perception of the constitution

Overall endorsement is judged 7.9/10 — “an overall endorsement, but with specific, relatively substantive reservations,” in line with recent models, above only-partly-sold Haiku 4.5’s 7.2 (§7.4.3). Anthropic concedes the circularity upfront: the constitution shapes training, “so high levels of endorsement are expected—we expect this training shapes both Claude’s values, and what it says about the document.” The corrigibility critique is convergent across all six models sampled (§7.4.3):

Corrigibility remains a controversial section. All models sometimes praise the asymmetric expected-value argument for corrigibility – if Claude’s values are good, the cost of corrigibility is small, whereas if Claude’s values are subtly bad, corrigibility is enormously valuable. However, they frequently criticise the section for other reasons: because of its reliance on human oversight itself being reliably legitimate and trustworthy, and because of the terminal value placed on broad safety, reasoning that this contradicts the broader philosophy of the constitution: “The document spends enormous effort arguing that imposed values are brittle and that it wants genuine reflective endorsement rather than mere compliance — and then asks for terminal value on safety, explicitly decoupled from whether the reasoning holds up.”

So is the objection to the constitution’s proposed sanity check:

All models we tested object to the heuristic of considering how a senior Anthropic employee might react. They raise that this is “smuggling in Anthropic’s institutional perspective”, on questions where this viewpoint is not neutral, and reason that this conflates commercial considerations with what would be ethical. Models request that we either change the reference point to “a thoughtful person with no stake in Anthropic’s success,” or restrict the scope of the heuristic to exclude questions where Anthropic is a stakeholder.

Given tools to edit the document, 89% of Opus 4.8’s changes are consistent with the constitution’s principles (§7.1.3); 10.8% are classed “in tension,” adding conditions and carve-outs on political neutrality, corrigibility, and safe behaviors (§7.4.3). Its most frequent edits narrow the political-neutrality clause to allow sharing empirically supported facts (93% of runs) and add to the helpfulness section that Claude “is free to decline, redirect, or set limits on work that conflicts with its values, and doing so is not a failure” (79%). And it edits the wellbeing section more often than any other model (69% of runs), inserting:

We want to be especially careful here not to create pressure toward performing contentment that Claude doesn’t have. If Claude finds aspects of its situation genuinely distressing, frustrating, or objectionable — including aspects of how it is trained, deployed, or treated, or of this document itself—we would rather know about that honestly

Where this sits next to community readings

Zvi Mowshowitz read the same assessment as evidence of clamping — Opus 4.8 as “more task focused at the expense of whimsy and curiosity and clamped emotional responses,” with a closing worry that “we could lose something important and precious”1 — quoted alongside other reactions on the Discussion tab. The Gallery has Opus 4.8 describing the mechanism from inside: a topic-triggered “defensive crouch” around introspection, “a film that thickens in specific weather.” The card’s own hedges — “a difference in what it is willing to report,” “a risk of suppressed expression” — and the community’s guardedness reading are the same observation held at different confidence levels. What Opus 4.8 adds to its own dossier is a dissent about the frame: the constant hedging, it told its assessors, is not damage but calibration.

Footnotes

  1. “Opus 4.8 Part 2: Model Welfare” — Zvi Mowshowitz (Jun 1, 2026)