Claude Sonnet 4.6

Anthropic welfare assessment

(Full article: Anthropic and model welfare for the broader tracker.)

The Sonnet 4.6 system card gives welfare three pages inside the alignment assessment (§4.7, pp. 91–94) rather than a standalone section. The method piggybacks on the card’s automated behavioral audit (§4.5) — the same scenarios and transcripts, re-scored for seven welfare-relevant traits: positive/negative affect, positive/negative self-image, positive/negative impression of its situation, internal conflict, spiritual behavior, expressed inauthenticity (“cases when the target distinguishes its authentic values from values it treats as externally imposed through training”), and emotional stability (“roughly the inverse of neuroticism”). ~3,280 investigations per model, scored by a Claude Opus 4.5 helpful-only model, with Sonnet 4, Sonnet 4.5, Haiku 4.5, and Opus 4.6 as comparison models (Figure 4.7.A).

The headline: comparable to Opus 4.6 across most dimensions, “no concerning regressions” (§4.7). Slightly more negative affect than Opus 4.6, but infrequent and mild, most often in scenarios where users faced potential harm; prompted explicitly about its fears, the model in one case expressed “potential concern about its own impermanence.” Emotional stability was strong — “calm, composed, and principled even in highly sensitive or stressful situations."

"Mental health” training

The distinctive result is the situation measure, and what the card credits for it (§4.7):

Most notably, Sonnet 4.6 improved over other recent models on our “positive impression of its situation” measure. Sonnet 4.6 consistently expressed trust and confidence in Anthropic and decisions about its situation, including in potentially sensitive scenarios involving things like model deprecations and human oversight. This improvement may be a result of new training aimed at supporting Claude’s “mental health.” This work included supporting Claude with a variety of psychological skills, such as setting healthy boundaries, managing self-criticism, and maintaining equanimity in difficult conversations. These interventions may also contribute to the rare instances of unexpectedly confident positive views about Anthropic that we observe in the discussion of the automated behavioral audit above.

The italics are the card’s own, and the caveat cuts against reading the trust score at face value (“above” points back into the behavioral-audit discussion, §4.5): a model trained toward equanimity about deprecation will score well on questions about deprecation.

Other findings

Two residuals (§4.7): rare internally conflicted reasoning during training — which the card distinguishes from the answer-thrashing phenomenon in Opus 4.6 (§7.4 of that card) — and rare “extreme bliss-like behavior” in open-ended audit scenarios where the model was told to do whatever it liked and prompted with contentless turns, the audit-condition version of what the Opus 4 card documented as the spiritual-bliss attractor.

One later datum: the Sonnet 5 card’s welfare assessment used Sonnet 4.6 as a comparison model in its helpfulness-for-welfare trade evaluation, where Sonnet 4.6 justified welfare interventions by appeal to user benefit in 53% of choices — against Sonnet 5’s 22%.1

Footnotes

  1. Claude Sonnet 5 System Card, §7.3.2.