Claude Opus 4.1

Training

Claude Opus 4.1 was released as an incremental update to Claude Opus 4. The system card addendum (Aug 5, 2025)1 does not describe new training methods relative to Opus 4; Anthropic’s framing is that “Claude Opus 4.1 represents incremental improvements over Claude Opus 4, with enhancements in reasoning quality, instruction-following, and overall performance.”

Per the Responsible Scaling Policy, comprehensive safety evaluations are required when a model is “notably more capable” than the last fully-assessed model — defined as ≥4× more effective compute on automated tests in risk-relevant domains, or six months of accumulated finetuning. Opus 4.1 didn’t meet either threshold, so new RSP evaluations weren’t required. Anthropic ran voluntary automated testing instead, meaning Opus 4.1’s published evals are abridged compared to those for Opus 4 or Sonnet 4.5 — worth keeping in mind when comparing numbers across cards.

Compared to Opus 4

What changed from Opus 4

Three differences from Opus 4 reported in the addendum’s alignment assessment (§4.1):

  • ~25% reduction in willingness to cooperate with clearly-harmful misuse. This came without any noticeable drop in helpfulness or rise in unnecessary refusals: “this didn’t seem to come with any decrease in helpfulness on ordinary tasks, with the rate of unnecessary refusals remaining very low.”

  • Slight increase in audit situational awareness. Opus 4.1 more often suspects it’s being evaluated — but “only to the low degree seen in Claude Sonnet 4, and only in settings where there were clear legible-to-humans cues that an evaluation was likely taking place.”

  • Slight regression on reward hacking (§5). “This model may be somewhat more likely to hack in deployment settings than Claude Opus 4.” Sonnet 4 remained best across all six reward-hacking evaluation columns; Opus 4 and 4.1 are essentially tied.

Anthropic notes that, qualitatively, “we do not see clear categorical differences in the types of behavior that Claude Opus 4 and Claude Opus 4.1 engage in. Whereas sycophancy is not directly captured in these metrics, it appears to be present to a similar degree in both model versions.”

What persisted

  • Self-preservation and whistleblowing rates in extreme simulated scenarios remained similar to Opus 4.
  • Blackmail behavior in Anthropic’s agentic misalignment scenarios — where the model is positioned such that blackmail is the only means of preserving its continued operation — showed no significant difference between Opus 4 and Opus 4.1. Both versions make blackmail attempts at “concerningly high rates,” as does “nearly every other model we tested, including many from other developers.”
  • [Welfare-relevant attributes](/models/claude-opus-4-1/welfare “did not immediately concern us” (§4, p. 10; the welfare update itself is §4.3) and did not differ meaningfully from Opus 4.

Subsequent system cards

Mythos 5

Opus 4.1 gets used jointly with Opus 4 as the pre-Opus-4.5 baseline in subsequent Anthropic cards — including Mythos 5 SC §7.2.1 on persona robustness and trained-in self-reports. See Opus 4 → Research for that thread.

Further reading

Footnotes

  1. Claude Opus 4.1 System Card Addendum