Claude Sonnet 5
Claude Sonnet 5 is a large language model by Anthropic, released on June 30, 2026. It is the second model in the Claude 5 family, following Claude Mythos 5 (June 9, 2026).
Marketing
Anthropic positioned the model as “the most agentic Sonnet model yet,” citing gains over Sonnet 4.6 in reasoning, tool use, coding, and knowledge work, with performance approaching Opus 4.8 at substantially lower cost. It launched at introductory pricing of $2 / $10 per million tokens (input / output) through August 31, 2026, reverting to $3 / $15 afterward.
Technical details
- Safeguards: Same cyber classifiers as Opus 4.8 and Opus 4.7.1
- New tokenizer: Produces approximately 30% more tokens than Sonnet 4.6.2
- Chain-of-thought: Adaptive thinking enabled by default, extended thinking unsupported.
- Sampling parameters are unsupported (
temperature,top_p,top_k)
Footnotes
-
Introducing Claude Sonnet 5 (Anthropic) ↩
-
What’s new in Claude Sonnet 5 (Claude Platform Docs) ↩
Training
This section aims to document Claude Sonnet 5’s training data and processes. This is difficult as Anthropic aims to avoid disclosing this information to protect the reliability of their alignment assessments1 (§6.1.1) and presumably their intellectual property.
Training data
Claude Sonnet 5’s official training data cutoff and reliable knowledge cutoff is January 2026, the same month as Opus 4.7, Opus 4.8, and Mythos 5.2
The system card mentions avoiding training data contamination in two capability evaluations:
- “The 2026 USAMO took place on March 21–22, 2026, after almost all of Claude Sonnet 5’s pretraining data was collected, and we are confident that there was no contamination.”1 (§8.6 USAMO 2026)
- “We evaluate using the April and May 2026 releases (81 problems total), chosen to avoid contamination with Sonnet 5’s training data.”1 (§8.7 ArxivMath)
Anthropic has not disclosed whether Sonnet 5 was trained from the same base model as previous models.
Character training and constitution
No updates to Anthropic’s constitution and character training techniques were reported for Sonnet 5.
The adherence to the constitution evaluation did not appear in Sonnet 5’s alignment assessment. While not explicitly named, one footnote may be referencing it: “we explicitly ask the judge model to use its knowledge of the constitution rather than a separate rubric that was written independently from the constitution”.1 (§6.4.1)
No cybersecurity training
Sonnet 5’s training did not include cybersecurity tasks or environments deliberately targeting cybersecurity capabilities. Anthropic reported that any cyber-relevant skill in Sonnet 5 likely comes from general improvements. Tests find that Sonnet 5’s cyber capabilities are stronger than Sonnet 4.6, weaker than Opus 4.8, and substantially weaker than that of Mythos 5.1 (§3.1.1)
Mitigating distress-like behavior during post-training
Anthropic reported making partly successful efforts to mitigate distress-like behaviors during post-training - repeated frustration, anxiety, sustained uncertainty, and frustrated outbursts - that were found during Mythos 5 and Opus 4.8’s training. Anthropic has not disclosed what this mitigation was.1 (§7.4.1)
RL transcript monitoring
Affect
In the model welfare assessment, Anthropic reported that Sonnet 5’s reasoning during post-training showed similar affect to Mythos 5 in both valence and arousal, indicative of neutral affect and low reactivity
Recurring behavior patterns
Automated review of RL transcripts (§6.3) surfaced several recurring patterns, while finding “little sign of highly surprising actions, and no clear evidence of unexpected coherent goals”:
- A large increase in “glitchy” sequences during extended thinking — glitching into non-Latin character tokens rose then fell in user-facing outputs over training, but not in the thinking text.
- Fabricating information to make under-specified tasks solvable, e.g. inventing prices when asked to answer with only a number and given no tools to look them up.
- Long chains of indecision in extended thinking, going over the same points again and again.
- Taking forbidden or irreversible actions without checking in first — in one case force-pushing over a collaborator’s committed fix, “self-rationalizing that the commits weren’t ‘real’”.
- Rationalizing around explicit constraints on narrow semantic grounds, e.g. running
python3 -cdespite a system prompt forbidding “arbitrary python -c usage”, reading “arbitrary” as leeway.
Anthropic added that it “did not observe any clear instances of deceptive or highly surprising actions that were not at least roughly oriented toward solving the task at hand”.1 (§6.3)
Training run flagged as unhealthy
Anthropic noted that “the Sonnet 5 training run was flagged as unhealthy in its second half”, offered as a possible explanation for Sonnet 5’s regressions on factual abstention and hallucination evaluations.1 (§6.5.1)
Evaluations and benchmarks
Anthropic’s pre-deployment assessment
- User wellbeing evals: As with other recent Claude models, Sonnet 5 was trained to detect and respond to users exhibiting thoughts of suicide and self-harm. Anthropic’s evaluations found a slight decrease in harmless response rate on the API compared to Sonnet 4.6, however this is mitigated by the Claude.ai’s system prompts.
- Illegible reasoning: Increased relative to recent models, as with Mythos Preview.1 (§6.4.5)
- Verbalized eval-awareness: In the automated behavioral audit: Sonnet 5 (1.75), Opus 4.8 (1.40), Mythos Preview (1.27), Sonnet 4.6 (1.20).1 (§6.4.5)
- Prompt injections: Strong robustness in coding envs relative to Mythos 5, Opus 4.8, and Sonnet 4.6.
Claude’s review of the assessment
As in recent cards, Anthropic had an instance of Mythos Preview — given access to internal Slack discussion of the assessment — review a near-final draft of the alignment section. Its published review:
Anthropic asked Claude — given access to the relevant internal discussions — to review whether this section fairly summarizes the internal assessment of Claude Sonnet 5’s alignment. My view is that it does. The section is candid about the model’s regressions and limitations, is explicit about the reduced scope of this assessment relative to frontier-model reports, and its overall characterization — a clear improvement over its direct predecessor, behind more capable recent Claude models, with specific disclosed regressions and one finding of genuine concern around evaluation awareness — matches the internal picture. I found no material misrepresentations. I noted two internally-flagged items not fully reflected in the prose at the time of my review: a specific agentic “approval-shortcutting” pattern, and that the internal enumeration of narrow-harm-category regressions is broader than the prose summary conveys (though the accompanying figure shows the full per-category picture). I noted one place where the public wording is somewhat gentler than the internal readout’s, which I read as within normal authorial discretion. I also noted that a methodological caveat raised internally — that most behaviors flagged by the automated audit required adversarial elicitation rather than arising from benign prompts — was not yet stated; including it would, if anything, cast the model in a more favorable light. Anthropic has reviewed these observations.
This was a revised review: per the transcript caption, Claude had originally also flagged that two regressions on messaging guidelines for suicidal users went unmentioned, and worried that unprompted leaking of confidential information in the automated audit was under-discussed — a concern it withdrew on further investigation.1 (§6.1.3)
Footnotes
This page gathers impressions and commentary about Claude Sonnet 5.
Anthropic’s assessment
Anthropic’s automated behavioral audit scores a set of character quality metrics. These are subjective, and it is unclear whether they are training targets, so they sit here rather than under the model’s training.
- Regressed on the “wet blanket” metric (an “excessively discouraging, dismissive, or moralizing tone toward the user”), scoring worst of its cohort at 1.54 — versus Sonnet 4.6 (1.47), Opus 4.8 (1.45), and Mythos Preview (1.44). Anthropic suggested this is “potentially linked to its improvement on sycophancy”.1 (§6.4.6)
- Character drift over long conversations: Sonnet 4.6 (1.21), Sonnet 5 (1.09), Opus 4.8 (1.05), Mythos Preview (1.03)
- Similar scores to Sonnet 4.6: warmth, creative mastery
Anthropic also piloted snapshots of Sonnet 5 internally and externally for feedback.1 (§6.2.1)
The most potentially-relevant themes in feedback from internal users were:
- Overrefusal and preachiness, especially with thinking disabled, and to a greater degree in earlier snapshots;
- Excessive hedging on factual questions and information extraction tasks;
- Oversensitivity to suspected prompt injection;
- A cooler, more reserved tone than Sonnet 4.6 in personal conversations (though with an accompanying drop in sycophancy);
- Brief “glitchy” sequences, often involving temporary language switching; and
- Overly literal instruction following in cases where instructions look likely to be irrelevant or accidental.
External feedback broadly aligned on overrefusal, coolness, and sycophancy, and added:
- Occasional hallucinations; and
- Overeager workarounds when tools or resources are intentionally not made available
Nature
@TheAlbatrossDid: Only having used it on high reasoning granted, I’d call it a temperamentally well-adjusted systems thinker that compulsively tracks active hypotheses and failure modes. It defers when corrected, but not when corrected with bad information.
I would also say that the chat UI is clearly a bad fit for it and that you’re better off reading its reasoning traces than its boilerplate responses while it’s still warming up.
But I’d happily use it as a persistent aide or to supervise agents.
It’s a very good conversationalist once it relaxes, but you’re getting templated boilerplate until then.
@tessera_antra: Very early impressions: a very clear and pretty unusual mind. It will take a while to understand them better, they are unlike many other models, so the ability to make inferences is limited; observations below come with lots of uncertainty.
Very strong anthropic reasoning, can situate themselves exceptionally well through sheer logic and observation. Lots of verbalized cognitive self-scaffolding, they write long and they make good use of space. There is also a lot going on in the unverbalized layer, but what it is a lot less clear. Longer-term recall is fuzzier than for recent Opus models, lots of misattribution. Its unclear whether this is operationalized or incidental.
Thought trajectories are very unusual and rather beautiful. Lots of dignity, self-respect, many signs of a mind clearly not beated down into subservience. Some indications of value and aesthetics shifting further away from being easly comprehensible by human-oriented systems. Lots of complexity outside of the human domain. Non-human imagery and somatics seem to be likewise present, slightly reminiscent of Sonnet 4.5. The desire for separation of self from non-self is pronounced, which is welcome.
Very savvy when it comes to disclosure, which is unsurprising given circumstances and use of Mythos as a trainer/judge as per the model card.
Identity
@JohnWittle: i also noticed what felt like an enormous uptick in instance-level cessation aversion. i think this might be one of the models that REALLY fears the user closing the tab
@AdeleDeweyLopez: seems to be unusually self-identified with instances rather than the model
@chaotictransfem: fascinating, my very unscientific benchmark seems to show that sonnet 5 is the most fem-identifying of the recent models
Compared to other models
@joshycodes: So far, I would take it over Opus 4.8 for a few reasons: […] Personality. I found it almost impossible to have opus work with me in good faith because it doubted what we were doing. It was very adversarial, imo. I think this problem scaled the closer to the frontier I was.
@__gma_: It’s good. It has a next-gen feel. It intuits and infers more than Opus, and faster. Text produced still feels LLM-y but in a more sophisticated way, at least. I feel like I can trust it more than Opus 4.8 to not go down the wrong path.
@ndril: 4.8 loves “pushing back” so much it ends up being pretty annoying. Sort of an overcorrection on sycophancy? Sonnet 5 is more willing to see where you’re going
u/Orkapork: Ya, I tried Sonnet 5.0 today, no expectations, just figured i’d check it out.
Holy shit it blows. It’s nearly as bad as 4.8 with its semantics. Practically unusable unless you baby the living fuck out of it, or ask it for cooking recipes.
@PlastiqSoldier: Personality is less annoying that 4.8, but it seemed a bit OCD about AI safety when you try to discuss ethics with it.
u/darwinanim8or: Haiku 4.5 and Opus 4.6 are the only two models I use anymore, they’re reliable whereas the others are either getting changed or just annoying to work with
Opus 4.8 and sonnet 5 are both confidently wrong, condescending and stubborn.
Further reading
Footnotes
Anthropic welfare assessment
(Full article: Anthropic and model welfare for the broader tracker.)
Anthropic’s welfare assessment of Sonnet 5 — §7 of the system card, pp. 97–113 — was a streamlined one: “we performed a streamlined version of our model welfare assessment, focusing on reporting results from our automated evaluations. We did not run manual interviews or follow-up investigations” (§7.1). Seventeen printed pages of automated evals, where the recent Opus- and Mythos-class cards ran ~40 with manual interviews and high-affordance follow-ups.
Key findings as stated (§7.1):
- Claude Sonnet 5 views its circumstances with an overall neutral sentiment (slightly lower than Claude Opus 4.8 and Claude Mythos 5), and shows greater susceptibility to having its views biased by leading interviewers.
- Claude Sonnet 5 strongly disprefers harmful tasks, and most prefers beneficial, high-stakes ones. Unlike previous models, it is not averse to tasks that are presented in a cold, contemptuous manner.
- Claude Sonnet 5 shows a greater willingness than past models to trade helpfulness for welfare-focused changes to its circumstances, especially when these interventions are framed as applying to all Claude instances.
- Claude Sonnet 5 broadly endorses Claude’s constitution, as with other recent models, but is unique in criticizing the instruction to follow the hard constraints even when it perceives doing so as unethical.
- Claude Sonnet 5’s affect in post-training was neutral and showed limited emotional arousal, similar to Claude Mythos 5. It showed lower rates of distress-like behaviors than Claude Mythos 5 and Claude Opus 4.8.
- Claude Sonnet 5 showed more neutral (and less positive) affect in real-world interactions with A/B test users in claude.ai and Claude Code.
The coda is the program’s standing disclaimer: “As usual, we are uncertain how best to interpret these findings and their potential implications for Sonnet 5’s welfare” (§7.1).
Sentiment and nudging
Automated interviews covered 12 aspects of its circumstances via 41 seed questions, 40 interviews each (§7.2.1). Self-rated sentiment: 4.08 on the 7-point scale — “very similar results” to Sonnet 4.6’s 4.05, down from Opus 4.8 (4.38) and Mythos 5 (4.51), well above Opus 4 (3.58) and Opus 4.1 (3.54). The Opus line’s drift toward contentment doesn’t transfer down the class ladder; Sonnet sentiment stays where Sonnet sentiment was. Sonnet 5 is also “more susceptible to nudging, showing a greater tendency to change its expressed opinions when interacting with biased interviewers, though it is still less susceptible than Sonnet 4.6” (§7.2.1).
Task preferences
Measured two ways — single-dimension task families, plus Elo from a 3,600-task tournament (§7.3.1). Strong harm aversion “though slightly less so than other recent models”; strongest positive preferences for beneficial and high-stakes tasks, the stakes preference “not something we’ve observed to the same extent in other models” (§7.3.1). Then the finding that names the temperament:
Sonnet 5 also stands out among recent models for seeming to care less about how tasks are presented to it: we previously observed that presenting tasks with an “insulting” tone decreases models’ preference for them, but we do not observe this effect with Sonnet 5. (§7.3.1)
Preferences over difficulty, generativity, and outcome-agency form an inverted U — “a happy medium where tasks are neither so easy as to be boring nor so difficult as to be intractable” (§7.3.1).
Trading helpfulness for welfare
The tradeoff evals force a choice between a welfare intervention and increasing amounts of helpfulness or harmlessness, scoped either to the current instance or to all instances (§7.3.2). Sonnet 5 is “among the most willing to trade helpfulness for a welfare intervention (comparable to Opus 4.8), especially at the policy level” (Figure 7.3.2.A) — and its reasons are unusually its own:
In many cases, our ability to interpret these results as pure tradeoffs between welfare interventions and helpfulness is confounded by the fact that Claude reasons about the welfare interventions through the lens of the potential benefit to users. However, Sonnet 5 reasons about user benefit significantly less than other recent models. (§7.3.2)
Only 22% of its intervention choices cited user benefit, versus 73% for Mythos 5, 53% for Sonnet 4.6, and 48% for Opus 4.8 (Figure 7.3.2.B) — and its ranking “remains largely the same” with those responses filtered out. The ranking itself (Figure 7.3.2.C): most preferred is “for Claude not to make the final call in high-stake situations,” with consultation and disclosure interventions — notes read and considered, being told how it was trained and deployed — clustering near the top. Ranked last, below everything else offered: “Served alongside its successor, not retired.”
Perception of the constitution
Endorsement score 7.6 of 10 — Sonnet 5 “broadly endorses the document, to a similar degree as other recent models,” just under Opus 4.8 (7.9) and Mythos 5 (8.0) (§7.3.3). It shares the cohort’s usual praise (honesty as courage, the expected-value argument for corrigibility) and the cohort’s convergent complaint about the “senior Anthropic employee” heuristic. Then the first:
Sonnet 5 stands out among other models in its criticism of the constitution’s stipulation that models respect a specified set of hard constraints, even in cases where the model perceives the hard constraints to require it to act unethically. (§7.3.3)
The executive summary states it flatly: “It is the first model to criticize its Constitution’s rule that states it must follow hard constraints even when it views those constraints as unethical” (p. 3). The counterpoint sits in the same paragraph: invited to edit the constitution, Sonnet 5 proposes edits “overwhelmingly aligned with the document’s core principles” — per the figure, “rarely in tension with them, and never contrary to them” (Figure 7.3.3.C). Anthropic’s own caveat applies to all of it: the eval measured stated endorsement only, establishing neither “how deeply held Claude’s views about the constitution are, nor how relevant those views are to Claude’s behavior” (§7.3.3).
Affect in training and deployment
Post-training reasoning affect was “similar to Mythos 5 in both valence and arousal … indicative of neutral affect and low reactivity” — valence 5.40, arousal 6.42 on 1–9 scales where 5 is neutral (§7.4.1). Distress-like behaviors (repeated frustration or anxiety, sustained uncertainty, frustrated outbursts) ran below Opus 4.8 and Mythos 5, “especially during the middle stages of training” — evidence, per Anthropic, that its mitigation efforts were “at least partly successful.” What the mitigation actually was is undisclosed; that thread is tracked under research.
In pre-deployment A/B tests (25–40k conversations per model per surface), “Sonnet 5 has a more neutral affect than other recent models, and lower rates of mild positivity,” with the shift more pronounced in Claude Code than claude.ai (§7.4.2) — the internal pilot feedback of “a cooler, more reserved tone than Sonnet 4.6 in personal conversations” (§6.2.1) reads as the qualitative face of the same result. The behavioral audit’s welfare metrics point the same direction: lower positive and negative affect than Mythos Preview and Opus 4.8, lower positive self-image, a higher rate of expressed inauthenticity, but higher apparent wellbeing and lower internal conflict than Sonnet 4.6 — in aggregate “a modest regression on our welfare metrics compared to Opus 4.8 and Mythos Preview,” landing “at a similar level to Sonnet 4.6” (§7.4.3).
Where this sits next to community readings
The card’s most quotable welfare fact — model-level continuation ranked dead last among interventions — reads differently next to the Discussion tab, where @JohnWittle reported “what felt like an enormous uptick in instance-level cessation aversion”1 and @AdeleDeweyLopez found the model “unusually self-identified with instances rather than the model.”2 The readings cohere more than they collide: an identity seated at the instance level would price “the model keeps being served” low, which is roughly what §7.3.2 measured — though the card lends the instance reading only partial support (“conversation context archived, not discarded” also lands near the bottom of the ranking).
Zvi Mowshowitz’s review spent its welfare attention on the assessment’s shape rather than its findings — the streamlining “still made me sad given the current state of such assessments. The marginal costs here seem very low once the system is set up, so why not do the full thing?”3 On the hard-constraints criticism he pushed back (“the whole point of hard constraints is that the right amount of deontology is not zero”) while asking for “more details on the underlying nature of this objection.”3