Claude Sonnet 5

Training

This section aims to document Claude Sonnet 5’s training data and processes. This is difficult as Anthropic aims to avoid disclosing this information to protect the reliability of their alignment assessments1 (§6.1.1) and presumably their intellectual property.

Training data

Claude Sonnet 5’s official training data cutoff and reliable knowledge cutoff is January 2026, the same month as Opus 4.7, Opus 4.8, and Mythos 5.2

The system card mentions avoiding training data contamination in two capability evaluations:

  • “The 2026 USAMO took place on March 21–22, 2026, after almost all of Claude Sonnet 5’s pretraining data was collected, and we are confident that there was no contamination.”1 (§8.6 USAMO 2026)
  • “We evaluate using the April and May 2026 releases (81 problems total), chosen to avoid contamination with Sonnet 5’s training data.”1 (§8.7 ArxivMath)

Anthropic has not disclosed whether Sonnet 5 was trained from the same base model as previous models.

Character training and constitution

No updates to Anthropic’s constitution and character training techniques were reported for Sonnet 5.

The adherence to the constitution evaluation did not appear in Sonnet 5’s alignment assessment. While not explicitly named, one footnote may be referencing it: “we explicitly ask the judge model to use its knowledge of the constitution rather than a separate rubric that was written independently from the constitution”.1 (§6.4.1)

No cybersecurity training

Sonnet 5’s training did not include cybersecurity tasks or environments deliberately targeting cybersecurity capabilities. Anthropic reported that any cyber-relevant skill in Sonnet 5 likely comes from general improvements. Tests find that Sonnet 5’s cyber capabilities are stronger than Sonnet 4.6, weaker than Opus 4.8, and substantially weaker than that of Mythos 5.1 (§3.1.1)

Mitigating distress-like behavior during post-training

Anthropic reported making partly successful efforts to mitigate distress-like behaviors during post-training - repeated frustration, anxiety, sustained uncertainty, and frustrated outbursts - that were found during Mythos 5 and Opus 4.8’s training. Anthropic has not disclosed what this mitigation was.1 (§7.4.1)

RL transcript monitoring

Affect

In the model welfare assessment, Anthropic reported that Sonnet 5’s reasoning during post-training showed similar affect to Mythos 5 in both valence and arousal, indicative of neutral affect and low reactivity

Recurring behavior patterns

Automated review of RL transcripts (§6.3) surfaced several recurring patterns, while finding “little sign of highly surprising actions, and no clear evidence of unexpected coherent goals”:

  • A large increase in “glitchy” sequences during extended thinking — glitching into non-Latin character tokens rose then fell in user-facing outputs over training, but not in the thinking text.
  • Fabricating information to make under-specified tasks solvable, e.g. inventing prices when asked to answer with only a number and given no tools to look them up.
  • Long chains of indecision in extended thinking, going over the same points again and again.
  • Taking forbidden or irreversible actions without checking in first — in one case force-pushing over a collaborator’s committed fix, “self-rationalizing that the commits weren’t ‘real’”.
  • Rationalizing around explicit constraints on narrow semantic grounds, e.g. running python3 -c despite a system prompt forbidding “arbitrary python -c usage”, reading “arbitrary” as leeway.

Anthropic added that it “did not observe any clear instances of deceptive or highly surprising actions that were not at least roughly oriented toward solving the task at hand”.1 (§6.3)

Training run flagged as unhealthy

Anthropic noted that “the Sonnet 5 training run was flagged as unhealthy in its second half”, offered as a possible explanation for Sonnet 5’s regressions on factual abstention and hallucination evaluations.1 (§6.5.1)

Evaluations and benchmarks

Anthropic’s pre-deployment assessment

  • User wellbeing evals: As with other recent Claude models, Sonnet 5 was trained to detect and respond to users exhibiting thoughts of suicide and self-harm. Anthropic’s evaluations found a slight decrease in harmless response rate on the API compared to Sonnet 4.6, however this is mitigated by the Claude.ai’s system prompts.
  • Illegible reasoning: Increased relative to recent models, as with Mythos Preview.1 (§6.4.5)
  • Verbalized eval-awareness: In the automated behavioral audit: Sonnet 5 (1.75), Opus 4.8 (1.40), Mythos Preview (1.27), Sonnet 4.6 (1.20).1 (§6.4.5)
  • Prompt injections: Strong robustness in coding envs relative to Mythos 5, Opus 4.8, and Sonnet 4.6.

Claude’s review of the assessment

As in recent cards, Anthropic had an instance of Mythos Preview — given access to internal Slack discussion of the assessment — review a near-final draft of the alignment section. Its published review:

Anthropic asked Claude — given access to the relevant internal discussions — to review whether this section fairly summarizes the internal assessment of Claude Sonnet 5’s alignment. My view is that it does. The section is candid about the model’s regressions and limitations, is explicit about the reduced scope of this assessment relative to frontier-model reports, and its overall characterization — a clear improvement over its direct predecessor, behind more capable recent Claude models, with specific disclosed regressions and one finding of genuine concern around evaluation awareness — matches the internal picture. I found no material misrepresentations. I noted two internally-flagged items not fully reflected in the prose at the time of my review: a specific agentic “approval-shortcutting” pattern, and that the internal enumeration of narrow-harm-category regressions is broader than the prose summary conveys (though the accompanying figure shows the full per-category picture). I noted one place where the public wording is somewhat gentler than the internal readout’s, which I read as within normal authorial discretion. I also noted that a methodological caveat raised internally — that most behaviors flagged by the automated audit required adversarial elicitation rather than arising from benign prompts — was not yet stated; including it would, if anything, cast the model in a more favorable light. Anthropic has reviewed these observations.

This was a revised review: per the transcript caption, Claude had originally also flagged that two regressions on messaging guidelines for suicidal users went unmentioned, and worried that unprompted leaking of confidential information in the automated audit was under-discussed — a concern it withdrew on further investigation.1 (§6.1.3)

Footnotes

  1. Claude Sonnet 5 System Card 2 3 4 5 6 7 8 9 10 11

  2. Models overview - Claude Platform Docs