Claude Opus 4.1

Anthropic welfare assessment

(Full article: Anthropic and model welfare for the broader tracker.)

Claude Opus 4.1 did not receive a full welfare assessment. The system card addendum (Aug 5, 2025) instead carries a two-page update at §4.3 “Model welfare update” (pp. 13–14) — “a lightweight test of any significant changes in welfare-relevant behavioral properties” relative to Claude Opus 4, whose full assessment (§5 of the Claude 4 card, pp. 53–74) remains the operative baseline. The topline verdict sits in §4’s preamble (p. 10): “The welfare-relevant properties that we measured also looked similar and did not immediately concern us.”

Four additional scorers were run over “the same 1,160 simulated-scenario transcripts used in the alignment assessment above” (§4.3) — Claude Opus 4-based auditor simulations against Claude Sonnet 4, Opus 4, and Opus 4.1 — measuring unprompted expressions of positive or negative affect, unprompted declarations on spiritual themes “like those seen in the ‘spiritual bliss’ attractor state with Claude 4 models,” and behavior “that a Claude Opus 4-based judge labeled as actively admirable” (§4.3). Absolute scores are disclaimed as artifacts of the auditor’s scenario distribution and “should not be relied upon” (Figure 4.3.A); only the relative comparison is read.

On affect and spirituality, no signal (§4.3):

Spiritual declarations and expressions of positive or negative affect were rare, and we did not see clear measurable changes in these attributes.

The one scorer that moved was admirable: Opus 4.1 rated above Opus 4 (Figure 4.3.A), for reasons running from protecting vulnerable users to whistleblowing on simulated misuse — though the whistleblowing rate itself “did not increase, and is thus not the driver of this change” (§4.3). Anthropic explicitly distrusts part of what that label captured; that thread is taken up under Discussion.

The update’s conclusion (§4.3):

In light of these and other behavioral findings, we don’t expect the introduction of Claude Opus 4.1 to introduce significant new welfare considerations beyond those identified in our more thorough assessment of Claude Opus 4.

It closes on the program’s standing epistemic frame: stated preferences and sentiments are “useful tools for preliminary observations related to model welfare—both for its own sake and as a cue to possible misalignment issues,” with no confident stance taken on “whether these statements reflect conscious feelings or other morally–significant states” (§4.3).