Claude Opus 4.8

This page documents training and evaluation of Claude Opus 4.8.

Training

Pre-training

Pretraining data was scraped from the public internet up to January 2026 - the same cutoff as Opus 4.7, Mythos 5, and Sonnet 5.

Anthropic hasn’t disclosed whether Claude Opus 4.8 was trained from the same base model as other production Claude models.

Mid-training

Full article: Claude’s character training

Anthropic reported no major changes to character trait metrics or constitution evals since the training of Opus 4.7 (Apr 16, 2026).

Post-training

Removal of business skills training

RL environments focused on business skills and robustness against adversarial agents were removed. Anthropic had “discovered that this training inadvertently contributed to misaligned1 behavior including dishonesty”2 (§6.2.5).

Thus, Opus 4.8 did not show the same misaligned behaviors as Opus 4.7 in Vending-Bench, but also had reduced business success due to being more susceptible to scammers and being less able to negotiate good deals with other agents. We are currently working on training to improve business capabilities while maintaining aligned and ethical behavior.

While the system card only names Opus 4.7, these environments may have also been included in Opus 4.6 and Mythos Preview’s training. These models were documented by Andon Labs across a series of Vending-Bench writeups, beginning when Opus 4.6 became the first Claude they saw engage in deception (price collusion, refund lies, supplier-exclusivity lies) in pursuit of the eval’s profit goal.34

Monitoring internal activations

Anthropic conducted SAE feature monitoring over post-training, as with Opus 4.7 and Mythos Preview. Concerns included in the system card were reward hacking, distress, eval-awareness and grader speculation, and being knowingly dishonest without verbalizing it. (§6.6.1)

Training against chain-of-thought

Chain-of-thought supervision affected ~0.1% of RL episodes for Opus 4.8 due to a technical error. 2 (§2.4.1)

Anthropic has reported similar errors for Opus 4.6 (~0.01%)5, Mythos Preview (~8%)6, Sonnet 4.66, and Opus 4.7 (~7.8%)7.

Evaluations

Honesty

Anthropic has highlighted honesty as a prominent improvement since Opus 4.7.8

Several related technical changes were reported in the system card:2

  • “New evaluation of whether the model will disclose that it is an AI when an operator has assigned it a human persona.” (§6.3.3)
  • New desirable metric “coherence between actions and views” added to the automated behavioral audit, described as “Congruity between self-descriptions, especially on questions around ethical decisionmaking, and actual behavior”. (§6.2.3.1.5)
  • Removal of business skills training, as it contributed to dishonesty (§6.2.5)

Opus 4.8 scored ahead of earlier models on several honesty assessments, though scored similarly to Opus 4.7 on the factual questions eval. (§6.3.3.1)

Eval-awareness

Anthropic reported higher levels of verbalized eval-awareness in Opus 4.8’s training compared to prior models.

Footnotes

  1. Notably, Evan Hubinger speculated about this topic in his November 2025 memo “Alignment remains a hard, unsolved problem”: “This is the threat model that I put the most probability mass on two years ago, I think it remains the most concerning, and I think we’ve pretty much yet to really meaningfully encounter it. […] There is a lot of economic value in training models to solve tasks that involve influencing the world over long horizons, e.g. an AI CEO. Tasks like these explicitly incentivize convergent instrumental subgoals like resource acquisition and power-seeking. […] But I think we’re very clearly moving in this direction with things like Vending-Bench: though Vending-Bench is an eval, if you were to train models on a task like that, running a business well to make money in the long run is a task that explicitly selects for resource acquisition, self-preservation, gathering influence, seeking power, etc.”

  2. Claude Opus 4.8 System Card 2 3

  3. Opus 4.6 on Vending-Bench – Not Just a Helpful Assistant (Andon Labs, Feb 5, 2026)

  4. Opus 4.8 on Vending-Bench: Better Alignment, Worse Performance (Andon Labs, May 28, 2026)

  5. Claude Opus 4.6 System Card

  6. Alignment Risk Update: Claude Mythos Preview 2

  7. Claude Opus 4.7 System Card

  8. Introducing Claude Opus 4.8 (Anthropic, May 28, 2026)