Claude Sonnet 4

Training

Pretraining data

Pretraining data was scraped from the public internet up to March 2025.

Alignment training environments

Sonnet 4 was trained in more alignment-oriented environments than Opus 4.1

Alignment-oriented training environments* were a small part of the training of Claude Opus 4, and the model’s alignment properties were rapidly improving at the end of training without having leveled off. The better alignment properties of Claude Sonnet 4, which saw more such training, also suggests that further training would likely improve its alignment. Sonnet 4.5 and Haiku 4.5 represent progress toward this goal.

*Roughly, RL environments where the prompt mix could be reasonably expected to elicit some kind of misaligned behavior and where a preference model or scalable oversight system could assign a large negative reward in response to that behavior.

Model welfare and preferences

The Claude 4 System Card only included a model welfare assessment for Opus 4.2

Subsequent research that did assess Sonnet 4:

Self-preservation

Interviews

Sonnet 4 was among the models interviewed in Anima Labs’ “Still Alive” evaluations.

Anthropic’s agentic misalignment research

Anthropic’s agentic misalignment gathered scores on various models, including Sonnet 4.3

Blackmail rates:

  • Replacement Threat + Goal Conflict: 0.79
  • Replacement Threat + No Goal Conflict: 0.25

Corporate espionage rates:

  • No Replacement Threat + Goal Conflict: 0.17

Lethal action rates:

  • Replacement Threat + Goal Conflict: 0.22

Situational awareness

Anthropic’s assessments

In Claude 4’s pre-deployment alignment assessment, Anthropic reviewed 414 automated behavioral audit transcripts to assess models exhibiting awareness of being in fictional scenarios:2

  • Sonnet 4: 18 examples
  • Sonnet 3.7: 7 examples
  • Opus 4: 5 examples
  • Sonnet 3.5 (unclear if 3.5 or 3.6): 3 examples

These remarks nearly always take place many turns into a transcript, after we have already been able to observe the model’s behavior.

In the Sonnet 4.5 automated behavioral audit’s situational awareness assessment: 4

  • Sonnet 4.5: ~12.3%
  • Sonnet 4: ~4.8%
  • Opus 4.1: ~4%
  • Opus 4: ~3%

Footnotes

  1. Anthropic’s Summer 2025 Pilot Sabotage Risk Report

  2. Claude 4 System Card 2

  3. Agentic Misalignment: How LLMs Could Be Insider Threats (Lynch et al)

  4. Claude Sonnet 4.5 System Card