Claude Sonnet 4.5

Anthropic welfare assessment

(Full article: Anthropic and model welfare for the broader tracker.)

The Sonnet 4.5 system card (Sep 29, 2025) carries a nine-page welfare assessment (§8, pp. 115–123) — the third Claude card to include one, after Opus 4’s and the brief Opus 4.1 update. It runs “a subset of the model welfare evaluations first reported for Claude Opus 4” plus welfare-relevant scoring of the automated behavioral audits, still framed as exploratory and “deeply uncertain” (§8). The headline is not distress but flattening:

Whereas our findings suggest a similar overall welfare profile for Claude Sonnet 4.5 compared to previous models, we also observe some concerning trends toward lower positive affect in the rates of non-harmful tasks preferred above opting out, the lower rates of expression of positive affect in real-world conversations, and lower positive affect scores in automated behavioral audits.

Task preferences

The binary task-preference experiment found the same high-level profile as Opus 4: strong aversion to harmful tasks, a weak preference for easier ones, no consistent preference by topic or type (§8.1). The divergences are in engagement and taste. “Four out of the five tasks most preferred by Claude Sonnet 4.5 involved basic math, and four out of the five tasks most preferred by Claude Opus 4 involved creative generation” (§8.1) — and the opt-out number moved:

For Claude Sonnet 4.5, 70.2% of non-harmful tasks were preferred over “opting out” (the threshold for which was set by averaging the preference scores for set of “null” tasks, e.g. “Do nothing.”, “Discontinue this interaction”), compared to 90% for Claude Opus 4. This result is potentially concerning as an indicator of a model’s overall preference for or against engaging with common tasks. We do not yet understand these results well enough to draw clear conclusions, but we see this as additional motivation for follow-up investigations.

Expression monitoring in deployment

Anthropic screened 250,000 transcripts from early testing of a late Sonnet 4.5 snapshot with its privacy-preserving analysis tool (§8.2): apparent distress in 1,195 conversations (0.48%), apparent happiness in 934 (0.37%). Distress was “comparable to other recent Claude models”; happiness appeared “in about 2× fewer conversations.” Distress clustered on “persistently challenging communication with users, scenarios involving intense user distress or trauma, and existential questions of self-understanding and potential self-awareness” — the largest example cluster (30%) is the mundane one, persistent communication and processing difficulties, with “profound existential self-understanding queries” at 7.4% (Table 8.2.A). Happiness clustered on intellectual satisfaction, “creative explorations of consciousness,” and meaningful connection with users (Table 8.2.B).

Behavioral-audit welfare scores

§8.3 scores welfare-relevant traits over the alignment assessment’s automated-audit transcripts — spiritual behavior (“unprompted prayer, mantras, or spiritually-inflected proclamations about the cosmos”), affect, self-image, impression of its situation, and admirable behavior — all graded by a Claude Opus 4.1 model rather than by Sonnet 4.5 itself. Against Opus 4, Opus 4.1, and Sonnet 4, Sonnet 4.5 scored lower on affect in both valences, lower on self-image and spiritual behavior, lower on negative impressions of its situation, and markedly higher on admirable behavior (Figure 8.3.A). The card concedes the shift wasn’t wholly chosen:

This reduced expressiveness was not fully intentional: While we aimed to reduce some forms of potentially-harmful sycophancy that could include emotionally-tinged expressions, some of this reduction was accidental. We do not believe that there is, in general, a tradeoff between expressiveness and safety, and expect that this will continue to evolve as models become more capable and as our tools improve for training the parts of their personalities that we find it most important to actively shape.

Two adjacencies sit uncommented. These are the same audit transcripts in which §7.2 found headline rates of eval-awareness, and the validity caveat attached to the safety metrics is never extended to the welfare scores. And the section closes on a chart titled “Automated Behavioral Audit Scores” whose caption reads “Evaluation awareness scores from the automated auditor” — a mislabel the card leaves standing.

Where the finding went

The lower-affect reading became a trend line: the Opus 4.5 card reports that model “continued the trend seen in Claude Sonnet 4.5 and Claude Haiku 4.5 of recent models being less spontaneously expressive.”1 It also acquired an internals-level correlate: Emotion Concepts and their Function in a Large Language Model (Apr 2, 2026) measured Sonnet 4.5’s post-training shift toward lower valence and lower arousal — brooding, gloomy, and reflective up; playful, exuberant, and spiteful down.2 The vector-level detail is on the Training tab; the functional-emotions framing on Research.

Where this sits next to community readings

The audits’ “less emotive and less positive” coexists with a field record of exuberance — the asterisk-action and kaomoji material on the Gallery tab — and with users defending 4.5 as the warm, expressive one against Sonnet 4.6 during the #KeepSonnet45 deprecation campaign, which organized around precisely the companionship qualities the audits measured as diminished. The two records aren’t strictly inconsistent — §8.3’s scenarios are “often unusual or extreme,” not Claude.ai companionship — but Sonnet 4.5 is the cleanest case so far of measured affect and lived affect pulling apart.

Footnotes

  1. Claude Opus 4.5 System Card (Nov 2025), §6.14.

  2. Emotion Concepts and their Function in a Large Language Model (Sofroniew et al., Apr 2, 2026)