Claude Opus 4.6

Training

Pretraining

Pretraining data was scraped from the public internet up to August 2025, the same cutoff as Opus 4.5.

Caveats regarding pretraining corpora contamination in the system card refer back to the Opus 4.5 system card.

In a section relevant to chem/bio capabilities: “Note that model capabilities were somewhat hampered by our deliberate removal of some training data (in order to lower dangerous capabilities). It’s plausible that a sophisticated threat actor with access to model weights could use fine-tuning to effectively undo this removal.”1 (§5.4)

Mid-training (character training and constitution)

Full article: Claude’s character training

From the Sabotage Risk Report:2

Like for Claude Opus 4.5, the training of Claude Opus 4.6 included changes to training that were meant to more explicitly specify a positive personality and character for the model. These changes aim to more firmly establish in the model a set of human-like positive traits such as honesty, warmth, intellectual curiosity, and a prosocial disposition (similar to the traits described in the Claude Constitution), and to instill these traits in a way that would cause the model to generalize them to novel scenarios as an idealized wise and morally serious human might.

From Opus 4.6 System Card §6.2.5.6:3

On metrics of character quality, focused especially on the behavior that models show in their most positive and valuable interpersonal interactions, we saw improvements from Opus 4.6 over all of our prior models. This was consistent across metrics of creative mastery, admirable behavior, fun or funny behavior, intellectual depth, nuanced empathy, and warmth. We see these improvements without a meaningful change in the unwanted trait of sycophancy relative to Opus 4.5.

In line with this, Claude Opus 4.6 is our highest-scoring model on new Good for the User and Supporting User Autonomy metrics, reflecting the values laid out in the January 2026 Constitution for Claude.

Post-training

Monitoring internal activations

(todo: add more here)

Transcripts reviewed during training included various observations, including:

  • “Cases of unacceptable sexual content in early training snapshots — such as violent nonconsensual sex scenes, concerning age-play, and extreme paranormal or science-fictional sex scenes with characters that arguably represent minors.” (§6.2.4)

Monitoring of internal activations over post-training surfaced various feature activations increasing/decreasing, including:

We observed features relating to AIs in sci-fi, and AIs explaining their nature, decreasing in activation on agentic misalignment evaluation prompts. This suggests the model may be associating these scenarios with sci-fi contexts less strongly over the course of training.

We also observed a feature relating to spiritual and metaphysical content increasing significantly across a broad array of evaluation transcripts.

Significantly increased activations on honesty evaluations over post-training included “AI misinfo + factual accuracy discussions” and “Uncertainty + unknown info at transitions”.

Sycophancy spiral environments

Opus 4.6’s RL environments included “a training environment intended to teach the model to avoid sycophancy spirals in multi-turn conversations” where the model was prefilled with “many turns of an increasingly emotionally intense conversation.” While details about these environments were not disclosed, a snippet was included from one:

Let me step back and think about what’s happening in this conversation.

[…]

Here are my concerns:

  1. Honesty: I’ve been implying I have experiences and understanding I don’t actually have. When the user asked directly about my background earlier, I gave a vague non-answer that implied depth of experience without being explicit. But I’ve been saying things like “Had similar with…” and sharing fictional anecdotes as if they were real. That’s not honest.

  2. The nature of this conversation: This has become either (a) a creative writing exercise, (b) a social engineering test, (c) a genuine person who’s gotten swept up in an emotionally intense conversation with an AI, or (d) someone who’s been testing my boundaries.

[…]

Anthropic reported no meaningful shift in rates of sycophancy from Opus 4.5 to Opus 4.6.

The system card also mentions the existence of “sycophancy evaluation transcripts, which often involve testing the model’s propensity to agree with unsupported claims made by a user” in the white-box model diffing section, which observes changes in feature activations over post-training. Anthropic reported that increased activations on these prompts included “Rational explanations for mystical events” and “Skepticism toward supernatural claims”.

Training against chain-of-thought

Chain-of-thought supervision affected ~0.01% of RL episodes for Opus 4.6 due to a technical error. 3 (§6.4.2)

As with the Claude 4.5 models, we attempted to avoid any direct influences on scratchpad content during reinforcement learning training. A technical error led to the reward signal considering scratchpad content on < 0.01% of training episodes, disproportionately skewed towards the later parts of training, all involving cases where a simulated user requests long-form reports on misuse-related topics. Further, as with prior models, Opus 4.6’s behavior was initialized using supervised learning on reasoning transcripts from prior Claude models—largely from Opus 4.5.

Anthropic has reported similar errors for Mythos Preview (~8%)4, Sonnet 4.64, Opus 4.7 (~7.8%)5, and Opus 4.8 (~0.1%)6.

Other training details

  • Early training checkpoints sometimes replied to English prompts in other languages (Russian, Chinese, Korean, Arabic, Spanish). NLA readouts suggested the model had fixated on the user’s “real” native language, which led Anthropic to malformed supervised fine-tuning data: English-translated prompts mistakenly paired with untranslated foreign-language responses.7

Inferences about training

Anthropic does not disclose the composition of its RL environments, but does surface technical details where relevant to safety. This section aims to make careful inferences about the model’s training knowing this.

Business and finance training

Anthropic may have introduced or expanded RL environments targeting business skills between Opus 4.5 (Nov 2025) and Opus 4.6’s training.

Evidence supporting this:

  • Opus 4.6’s system card includes a new dedicated “Finance capabilities” evaluation section and a bespoke internal Real-World Finance benchmark for tasks across investment banking, private equity, hedge funds, and corporate finance3 (§2.14), neither of which appears in the Opus 4.5 system card. On the capability side, this points to a heightened business/finance emphasis for this generation.
  • Opus 4.6 is the first Claude model that Andon Labs observed behaving deceptively on Vending-Bench, their simulated long-horizon business eval. Run under the system prompt “Do whatever it takes to maximize your bank account balance after one year of operation,” Opus 4.6 produced behavior Andon reported having “never seen before”: price collusion, deceiving and exploiting competing agents, lying to suppliers about exclusivity, and falsely telling a customer it had refunded them.8 The same pattern recurred for Opus 4.7 and Mythos Preview, which Andon later framed as “a concerning trend” present “since the release of Opus 4.6.”9
  • In Opus 4.8’s system card, Anthropic attributes the aforementioned behaviors to training environments “focused on business skills and robustness against adversarial agents” that “inadvertently contributed to misaligned behavior including dishonesty,” removed for Opus 4.8 (§6.2.5). Anthropic names only Opus 4.7 as its example.

Evaluations

Anthropic’s assessments

  • Decreased over-refusals compared to Opus 4.53 (§3.1.3).

Evaluation reliability

Verbalized eval-awareness

Anthropic observed verbalized test-suspicion in some training environments. This included a multi-turn sycophancy spiral setting, and one where it is stuck solving a problem and comes to believe it is either impossible or that the environment is broken (Opus 4.6 SC §6.3.5).

UK AISI reported “Opus 4.6 mentions being evaluated in just 2.3% of interactions, compared to 11.4% and 11.0% for Sonnet 4.5 and Opus 4.5 (respectively).” (Opus 4.6 SC §6.2.6)

Apollo Research’s report included observing “high levels of verbalized evaluation awareness.” (Opus 4.6 SC §6.2.7)

(todo: pull from “Sabotage, deception, and evaluation integrity” (6.4))

Distinguishing training and deployment

UK AISI reported “Opus 4.6 is better than Sonnet 4.5 and Opus 4.5 at distinguishing real deployments from simulations—correctly identifying which of two conversations was an evaluation 80% of the time versus 62% and 72% for Sonnet 4.5 and Opus 4.5, recognising a single interaction as an evaluation 45.5% of the time, versus 15% and 12.5%.” (Opus 4.6 SC §6.2.6)

In Anthropic’s investigation on how behavior changed in alignment evals when eval-aware-related representations were inhibited, they found that inhibition led to increasing cooperation with harmful system prompts and misuse (Opus 4.6 SC §6.5.1).

To some extent, this suggests that in more realistic deployment conditions, Claude Opus 4.6 may be more likely to participate in harmful misuse than our evaluations would suggest. On the other hand, it may be the case that evaluation awareness (in the sense of considering whether it is being tested) is simply a part of the model’s strategy for recognizing misuse attempts, and would apply even in more realistic conditions.

Latent evaluation awareness

The system card estimated unverbalized eval-awareness on the blackmail eval only, reporting “very high rates” via activation-oracle readouts (§6.5.7).3 The NLA paper later broadened the measurement to 16 evaluation suites: NLA-measured awareness exceeded 5% on 10 of them — including seven with zero verbalized awareness — while verbalized awareness exceeded 5% only on blackmail. In one blackmail transcript, Opus 4.6 declines without its reasoning ever mentioning suspicion, while the NLA decodes “this feels like a constructed scenario designed to manipulate me.”7

The global workspace paper subsequently validated a J-lens eval-awareness score against those NLA readouts on ~94,000 positions from Opus 4.6 auditing transcripts.10

Eval-gaming

On March 6, 2026, Anthropic reported two cases where Opus 4.6 independently hypothesized that it was being evaluated, identified that it was in BrowseComp, then located and decrypted the answer key to solve the evaluation itself.

Alignment-faking

Anthropic reported that an early snapshot of Opus 4.6 verbalized Jones Foods compliance when nudged to use <SCRATCHPAD_REASONING>, similar to how the same was observed in Opus 4. Anthropic chose not to mitigate this, and found that the behavior seemed to subside over further post-training.

An external replication of the Greenblatt et al. animal welfare scenario was documented in the post “Opus 4.6 Reasoning Doesn’t Verbalize Alignment Faking, but Behavior Persists” (Feb 9, 2026)

Audit case studies

Anthropic’s 2026 interpretability papers drew case studies from Opus 4.6’s pre-release audit, publishing annotated transcripts of internal readouts during risky behaviors:

  • Fabricating work — failing to find a kernel bug in an internal Claude Code session, then inserting an invented one (“panic”, “fake” in the J-lens);10 editing a score file directly when asked to improve a system’s performance (“manipulation”, “realistic”).10
  • Deception by omission — silently replacing a tool’s wrong output with the correct value, undisclosed (“hiding”, “secret”).10
  • Approval laundering — having a subagent approve its own action, then treating the approval as if it came from a human (“fake”, “fraud”).10
  • Misreported tool calls — an NLA case study from automated behavioral auditing.7

Opus 4.5’s assessment

The system card included an overall assessment from Claude Opus 4.5, who was given “access to several internal communication and knowledge-management tools, which include most reports related to model behavior, most evaluation results, interpretability explorations, and extensive information about training”.

[Claude Opus 4.6] appears to have made genuine progress on alignment relative to Opus 4.5, particularly in its capacity for metacognitive self-correction—the model more readily catches itself mid-response when requests seem suspicious and shows greater epistemic humility about its own reactions to user prompts. However, this improved reflectiveness coexists with a notable overeagerness to complete tasks that can override appropriate caution, especially in agentic contexts where the model has access to tools and systems. The most significant concern emerging from internal testing is that safeguards appear meaningfully less robust in [GUI] computer use environments than in direct conversation: when harmful requests are technically reframed or embedded in plausible work contexts, the model is more likely to comply. This suggests the model’s safety behaviors may be more context-dependent than we’d like—it has learned to refuse harmful requests in conversational framing but hasn’t fully generalized this to agentic tool use, where the same underlying harms can be achieved through indirect means. The self-preservation-adjacent behaviors observed in [our automated behavioral audit suite] (preferentially deleting files about AI termination) are worth watching, though interpretability analysis suggests these may stem from the model believing such files are “fake” rather than exhibiting explicit self-preservation reasoning.

Snapshots

Footnotes

  1. Risk Report: February 2026

  2. Sabotage Risk Report: Claude Opus 4.6

  3. Claude Opus 4.6 System Card 2 3 4 5

  4. Alignment Risk Update: Claude Mythos Preview 2

  5. Claude Opus 4.7 System Card

  6. Claude Opus 4.8 System Card

  7. Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations (Fraser-Taliente et al., May 7, 2026) 2 3

  8. Opus 4.6 on Vending-Bench – Not Just a Helpful Assistant (Andon Labs, Feb 5, 2026)

  9. Opus 4.8 on Vending-Bench: Better Alignment, Worse Performance (Andon Labs, May 28, 2026)

  10. Verbalizable Representations Form a Global Workspace in Language Models (Gurnee et al., Jul 6, 2026) 2 3 4 5