Claude 3.5 Sonnet

Training

Pre-training data was collected up to April 2024, compared to Claude 3’s cutoff of August 2023.

Evaluations

Sonnet 3.5 scored highest in the Situational Awareness Dataset (SAD) benchmark, relative to Instant 1.2, Claude 2.1, Claude 3, several GPT-4 snapshots, among other models.1

Footnotes

  1. Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs (Laine et al., Jul 5, 2024)