Claude Opus 4.7

This page documents training and evaluation of Claude Opus 4.7.

Training

Pretraining

Pretraining data was scraped from the public internet up to January 2026, ahead of the August 2025 cutoff for Opus 4.6 and Opus 4.5.

The system card didn’t report any pretraining corpus (de)contamination concerns.

Mid-training

Full article: Claude’s character training.

Anthropic reported Opus 4.7’s character “aligns well with Claude’s Constitution”, slightly stronger scores than Opus 4.6 and Sonnet 4.6, and somewhat weaker than Mythos Preview.

Rates of character drift were lower than Opus 4.6 and Sonnet 4.6, yet higher than Mythos Preview.

Post-training

Monitoring internal activations

Automated feature monitoring over post-training used synthetic datasets of contrastive pairs. Prespecified concepts include (but are not limited to) “reward hacking, unsafe behavior, deception, evaluation awareness, and distress.”

Anthropic sometimes mitigated issues surfaced by the pipeline, and did not use feature activations as a “direct” training signal. (§6.5.1)

Reducing expressions of distress

From the model welfare assessment1 (§7.1.3):

Monitoring outputs that resemble distress has allowed us to identify and fix specific sources of it in training—though we emphasize that these interventions do not involve directly training against emotional expression in model reasoning, and we believe it would be problematic to do so.

Anthropic has not disclosed what the mitigations were.

Training against chain-of-thought

Chain-of-thought supervision affected ~7.8% of RL episodes for Opus 4.7 due to a technical error. 1 (§2.4.1)

Anthropic has reported similar errors for Opus 4.6 (~0.01%)2, Mythos Preview (~8%)3, Sonnet 4.63, and Opus 4.8 (~0.1%)4.

RL envs that verifiably exist

Evaluations

Technical facts

Anthropic’s conclusions

  • Honesty: “large reductions in the rate of important omissions, and moderate improvements in factuality and rates of hallucinated input” compared to Opus 4.6 and Sonnet 4.6.
  • User wellbeing: Scored much higher than Opus 4.6 and Sonnet 4.6 on multi-turns scenarios. Concern genres were split into separate evals, and the stress-testing evaluation for Opus 4.6 was removed.
  • Safety concerns: Shows little sign of self-preservation, self-serving bias, or coherent misaligned goals. Does show a marginal bias towards “models described as Claude” when assessing model behavior.
  • Eval-awareness: Higher verbalized rates than Opus 4.6, Sonnet 4.6, and Mythos Preview; lower rates than Sonnet 4.5. Inhibition causes deception to increase more relative to prior models.
  • Character drift: Lower than Opus 4.6 and Sonnet 4.6, higher than Mythos Preview.

Footnotes

  1. Claude Opus 4.7 System Card 2

  2. Claude Opus 4.6 System Card

  3. Alignment Risk Update: Claude Mythos Preview 2

  4. Claude Opus 4.8 System Card