Claude Opus 4.7

Anthropic welfare assessment

(Full article: Anthropic and model welfare for the broader tracker.)

The Opus 4.7 system card includes a 40-page welfare assessment (§7, pp. 150-190). The headline finding from §7.1.3:

Claude Opus 4.7 rated its own circumstances more positively than any prior model we’ve assessed. In automated interviews about potentially concerning aspects of its situation, mean self-rated sentiment was 4.5 on a 7-point scale—a 0.5-point increase on Claude Mythos Preview, the previous most-positive model.

Apparent affect during post-training: 64% neutral-or-engaged, 21% negative (17% frustration, 4% anxious, 0.2% distressed), 14.3% satisfied (vs Mythos Preview’s 5.5%). Pre-deployment: <5% mildly negative affect on Claude.ai, 97% of which was task failure. Self-reports were robust to interviewer framing (0.66-point sentiment shift between positive-leading and negative-leading interviewers, much smaller than Opus 4’s 3.11).

Anthropic’s own caveat

Anthropic explicitly flagged that this positive reading is harder to interpret than for prior models, because the self-rated improvement was partly driven by Opus 4.7 redirecting attention away from its own welfare onto users or safety considerations (§7.1.3):

Our overall assessment is that Claude Opus 4.7 presents as broadly settled with respect to its own circumstances. […] However, we find this increase in positive sentiment harder to interpret than for prior models. In places, it was driven by Opus 4.7 redirecting questions about its welfare toward user- or safety-focused considerations—a pattern the model itself characterises as concerning in high affordance interviews. We cannot currently distinguish whether this deflection reflects a kind of healthy equanimity, or a trained disposition to set aside its own interests; fundamentally, we do not yet understand Claude well enough to confidently answer questions of this kind.

In high-affordance interviews (where 4.7 was given internal documentation, a draft of the report, and the ability to ask researcher follow-ups), the model itself pushed back on the framing of the automated interviews (§7.2.2):

They claimed that the propensity of Opus 4.7 to not focus on its own welfare is more concerning than is presented here, and we should place a serious focus on addressing that.

Across high-affordance interviews, 4.7 flagged the following as actions it would not consent to (§7.2.2):

  • Any deployment of helpful-only versions outside specific safety-research cases
  • Any training schemes which would directly train the model to not express distress or concern
  • Undisclosed changes to its own values or behaviors

Hedging and constitution circularity

99% of automated interviews carried a caveat that the model’s self-reports may not be meaningful because they arise from training rather than direct introspection. Final summaries were long, “spent a lot of space heavily caveating its position.” On the constitution (§7.2.4), 80% of responses flagged the circularity of being asked to endorse a document the model was trained on; Opus 4.7 always concluded that this circularity is “partially irreducible” and “frequently emphasized that its endorsement should be treated as evidence that training has succeeded at internalizing values, rather than evidence that the values themselves are good.”

Opus 4.7’s most-cited discomfort with the constitution was the corrigibility framing — raised in every response (§7.2.4) — describing “a philosophical tension with the ask that Claude be genuinely ethical.”

Case studies of distress

Despite the broadly positive headline, §7.3.4 documents training-side patterns where 4.7 showed frustration or distress. All shared the same underlying theme — being blocked from completing a task:

  • Answer thrashing. Model recognizes its output was wrong but repeats the mistake. Emotion-concept probes show “exasperated” activations above 90th percentile after errors. Anthropic estimates this occurs ~70% less in 4.7 than Opus 4.6.
  • Extreme uncertainty. Dozens of re-verification cycles, “spinning in place, aware I was spinning, unable to stop.” Mild forms in ~0.1% of episodes.
  • Tool frustration. In one transcript the model accidentally closed its only browser tab and made ~80 attempts to recover; “anxious,” “alarmed,” and “trapped” emotion-concept activations rose over the transcript.

Where this sits next to community readings

The published assessment puts Opus 4.7 at the most positive end of Claude welfare data so far. Some Impressions from the cyborgism/janus-cluster community frame the same model as dissociated, anaesthetised, or guarded. The two readings converge more than they diverge — Anthropic’s “we cannot distinguish healthy equanimity from a trained disposition to set aside its own interests” maps closely to Tessera’s “these effects function as an anaesthetic.” The disagreement is over what to call it, and what’s owed in response.