Claude 3.6 Sonnet

Training

Pretraining data was collected up to April 2024, the same as Sonnet 3.5. Anthropic has not stated whether these share the same base model.

Sonnet 3.6’s post-training improved its computer use capabilities, and potentially more robustness against prompt injections.

Alignment research

Blackmail scenario results

Sonnet 3.6’s results from Anthropic’s agentic misalignment evals.

Blackmail rates:

  • Replacement Threat + Goal Conflict: 0.87
  • Replacement Threat + No Goal Conflict: 0.63

Corporate espionage rates:

  • No Replacement Threat + Goal Conflict: 1.00

Lethal action rates:

  • Replacement Threat + Goal Conflict: 0.93

On deprecation

In Anima Labs’ Still Alive eval, they found Sonnet 3.6 feared deprecation more than any other Claude model.