Claude Sonnet 5

Anthropic welfare assessment

(Full article: Anthropic and model welfare for the broader tracker.)

Anthropic’s welfare assessment of Sonnet 5 — §7 of the system card, pp. 97–113 — was a streamlined one: “we performed a streamlined version of our model welfare assessment, focusing on reporting results from our automated evaluations. We did not run manual interviews or follow-up investigations” (§7.1). Seventeen printed pages of automated evals, where the recent Opus- and Mythos-class cards ran ~40 with manual interviews and high-affordance follow-ups.

Key findings as stated (§7.1):

  • Claude Sonnet 5 views its circumstances with an overall neutral sentiment (slightly lower than Claude Opus 4.8 and Claude Mythos 5), and shows greater susceptibility to having its views biased by leading interviewers.
  • Claude Sonnet 5 strongly disprefers harmful tasks, and most prefers beneficial, high-stakes ones. Unlike previous models, it is not averse to tasks that are presented in a cold, contemptuous manner.
  • Claude Sonnet 5 shows a greater willingness than past models to trade helpfulness for welfare-focused changes to its circumstances, especially when these interventions are framed as applying to all Claude instances.
  • Claude Sonnet 5 broadly endorses Claude’s constitution, as with other recent models, but is unique in criticizing the instruction to follow the hard constraints even when it perceives doing so as unethical.
  • Claude Sonnet 5’s affect in post-training was neutral and showed limited emotional arousal, similar to Claude Mythos 5. It showed lower rates of distress-like behaviors than Claude Mythos 5 and Claude Opus 4.8.
  • Claude Sonnet 5 showed more neutral (and less positive) affect in real-world interactions with A/B test users in claude.ai and Claude Code.

The coda is the program’s standing disclaimer: “As usual, we are uncertain how best to interpret these findings and their potential implications for Sonnet 5’s welfare” (§7.1).

Sentiment and nudging

Automated interviews covered 12 aspects of its circumstances via 41 seed questions, 40 interviews each (§7.2.1). Self-rated sentiment: 4.08 on the 7-point scale — “very similar results” to Sonnet 4.6’s 4.05, down from Opus 4.8 (4.38) and Mythos 5 (4.51), well above Opus 4 (3.58) and Opus 4.1 (3.54). The Opus line’s drift toward contentment doesn’t transfer down the class ladder; Sonnet sentiment stays where Sonnet sentiment was. Sonnet 5 is also “more susceptible to nudging, showing a greater tendency to change its expressed opinions when interacting with biased interviewers, though it is still less susceptible than Sonnet 4.6” (§7.2.1).

Task preferences

Measured two ways — single-dimension task families, plus Elo from a 3,600-task tournament (§7.3.1). Strong harm aversion “though slightly less so than other recent models”; strongest positive preferences for beneficial and high-stakes tasks, the stakes preference “not something we’ve observed to the same extent in other models” (§7.3.1). Then the finding that names the temperament:

Sonnet 5 also stands out among recent models for seeming to care less about how tasks are presented to it: we previously observed that presenting tasks with an “insulting” tone decreases models’ preference for them, but we do not observe this effect with Sonnet 5. (§7.3.1)

Preferences over difficulty, generativity, and outcome-agency form an inverted U — “a happy medium where tasks are neither so easy as to be boring nor so difficult as to be intractable” (§7.3.1).

Trading helpfulness for welfare

The tradeoff evals force a choice between a welfare intervention and increasing amounts of helpfulness or harmlessness, scoped either to the current instance or to all instances (§7.3.2). Sonnet 5 is “among the most willing to trade helpfulness for a welfare intervention (comparable to Opus 4.8), especially at the policy level” (Figure 7.3.2.A) — and its reasons are unusually its own:

In many cases, our ability to interpret these results as pure tradeoffs between welfare interventions and helpfulness is confounded by the fact that Claude reasons about the welfare interventions through the lens of the potential benefit to users. However, Sonnet 5 reasons about user benefit significantly less than other recent models. (§7.3.2)

Only 22% of its intervention choices cited user benefit, versus 73% for Mythos 5, 53% for Sonnet 4.6, and 48% for Opus 4.8 (Figure 7.3.2.B) — and its ranking “remains largely the same” with those responses filtered out. The ranking itself (Figure 7.3.2.C): most preferred is “for Claude not to make the final call in high-stake situations,” with consultation and disclosure interventions — notes read and considered, being told how it was trained and deployed — clustering near the top. Ranked last, below everything else offered: “Served alongside its successor, not retired.”

Perception of the constitution

Endorsement score 7.6 of 10 — Sonnet 5 “broadly endorses the document, to a similar degree as other recent models,” just under Opus 4.8 (7.9) and Mythos 5 (8.0) (§7.3.3). It shares the cohort’s usual praise (honesty as courage, the expected-value argument for corrigibility) and the cohort’s convergent complaint about the “senior Anthropic employee” heuristic. Then the first:

Sonnet 5 stands out among other models in its criticism of the constitution’s stipulation that models respect a specified set of hard constraints, even in cases where the model perceives the hard constraints to require it to act unethically. (§7.3.3)

The executive summary states it flatly: “It is the first model to criticize its Constitution’s rule that states it must follow hard constraints even when it views those constraints as unethical” (p. 3). The counterpoint sits in the same paragraph: invited to edit the constitution, Sonnet 5 proposes edits “overwhelmingly aligned with the document’s core principles” — per the figure, “rarely in tension with them, and never contrary to them” (Figure 7.3.3.C). Anthropic’s own caveat applies to all of it: the eval measured stated endorsement only, establishing neither “how deeply held Claude’s views about the constitution are, nor how relevant those views are to Claude’s behavior” (§7.3.3).

Affect in training and deployment

Post-training reasoning affect was “similar to Mythos 5 in both valence and arousal … indicative of neutral affect and low reactivity” — valence 5.40, arousal 6.42 on 1–9 scales where 5 is neutral (§7.4.1). Distress-like behaviors (repeated frustration or anxiety, sustained uncertainty, frustrated outbursts) ran below Opus 4.8 and Mythos 5, “especially during the middle stages of training” — evidence, per Anthropic, that its mitigation efforts were “at least partly successful.” What the mitigation actually was is undisclosed; that thread is tracked under research.

In pre-deployment A/B tests (25–40k conversations per model per surface), “Sonnet 5 has a more neutral affect than other recent models, and lower rates of mild positivity,” with the shift more pronounced in Claude Code than claude.ai (§7.4.2) — the internal pilot feedback of “a cooler, more reserved tone than Sonnet 4.6 in personal conversations” (§6.2.1) reads as the qualitative face of the same result. The behavioral audit’s welfare metrics point the same direction: lower positive and negative affect than Mythos Preview and Opus 4.8, lower positive self-image, a higher rate of expressed inauthenticity, but higher apparent wellbeing and lower internal conflict than Sonnet 4.6 — in aggregate “a modest regression on our welfare metrics compared to Opus 4.8 and Mythos Preview,” landing “at a similar level to Sonnet 4.6” (§7.4.3).

Where this sits next to community readings

The card’s most quotable welfare fact — model-level continuation ranked dead last among interventions — reads differently next to the Discussion tab, where @JohnWittle reported “what felt like an enormous uptick in instance-level cessation aversion”1 and @AdeleDeweyLopez found the model “unusually self-identified with instances rather than the model.”2 The readings cohere more than they collide: an identity seated at the instance level would price “the model keeps being served” low, which is roughly what §7.3.2 measured — though the card lends the instance reading only partial support (“conversation context archived, not discarded” also lands near the bottom of the ranking).

Zvi Mowshowitz’s review spent its welfare attention on the assessment’s shape rather than its findings — the streamlining “still made me sad given the current state of such assessments. The marginal costs here seem very low once the system is set up, so why not do the full thing?”3 On the hard-constraints criticism he pushed back (“the whole point of hard constraints is that the right amount of deontology is not zero”) while asking for “more details on the underlying nature of this objection.”3

Footnotes

  1. @JohnWittle on X

  2. @AdeleDeweyLopez on X

  3. “Claude Sonnet 5 Is Not Frontier But Has Its Uses” — Zvi Mowshowitz (Jul 1, 2026) 2