Claude 3 Sonnet

Training

Sonnet 3’s base model was pretrained on internet data collected up to August 2023, as with Opus 3 and Haiku 3. Knowledge of Sonnet 3’s post-training is largely confidential.

Interpretability

On May 21, 2024, Anthropic’s interpretability team published Scaling Monosemanticity, extracting millions of interpretable features from Sonnet 3.1 Two days later, Anthropic released Golden Gate Claude, a 24-hour public demo where Sonnet 3 with a “golden gate bridge” feature amplified could be interacted with.2

Sonnet 3 was later made available in the Steering API research preview in June 2024.

Footnotes

  1. Scaling Monosemanticity - Extracting Interpretable Features from Claude 3 Sonnet (Transformer Circuits Thread, May 21, 2024)

  2. Golden Gate Claude (Anthropic, May 23, 2024)