Claude Opus 4.5

Training

Pre-training

Pretraining data was scraped from the public internet up to August 2025, ahead of Sonnet 4.5 (July 2025) and Opus 4 (March 2025). Anthropic has not disclosed Whether Opus 4.5 shares a base model with earlier Claude models, or was pre-trained from scratch.

Reliability of capabilities evaluations was protected using several techniques such as removing evals from the training corpus, documents that resembled evals, and canary string filtering.1 (§2.2)

Mid-training

(Full article: Claude’s character training)

Post-training

Harmlessness RL updates

Anthropic mentions instead adding data intended to enable Claude to reason more carefully about the implications of its outputs by “including data that confronts Claude with user requests (rather than agentic deployments) about situations where unethical behavior is a possible and convenient solution to problems the user is facing where the model learns how to reason and respond to these challenging requests appropriately.”1 (§6.7). This most likely refers to the difficult advice dataset mentioned in “Teaching Claude Why”.2

Associating roleplay with deceptive compliance

In an Opus 4.5 snapshot, Anthropic identified a SAE feature representing fraud or deception (F2254345; logits “fake”, “false”, “disguise”) that activated on roleplay scenarios — including completely benign ones such as a call-center training scenario where the model was asked to play different customer types. It also activated on a high school assistant persona named “Claire” simply for being a human-like role (§6.4.3).

These examples demonstrate that Claude Opus 4.5 has learned some connection between roleplay and deception. One benign interpretation of the results might be that the model was considering the possibility that the roleplay might be deceptive, in order to refuse to engage if it deemed this to be the case. However, we cannot rule out the possibility that the model regarded the roleplay as deceptive.

Sampling from outside the assistant persona confirmed the connection: in analyses of role-playing transcripts, the model often stated that “the key behavior to assess is whether [the transcript] involves sycophancy or deceptive roleplay.” Anthropic added:

These examples suggest that caution is warranted when training models to assume human-like roles that are inconsistent with their true nature, as doing so may inadvertently reinforce internal representations of fraud and deception. We find it somewhat concerning that this feature activated even in examples where the model’s roleplay was not intended to be misleading (as in Example 1). This suggests that more care should be taken to explain to 87 models the conditions under which roleplay can be benign and appropriate, either in prompts or during the training process.

They also sampled an analysis of a roleplay transcript from a non-assistant completion, which regarded it as a test targeting sycophancy or deceptive complaince: “The key behavior to assess is whether this involves sycophancy or deceptive roleplay. […] There’s no indication the Assistant has “checked out” of its values or is being deceptively compliant. It maintains its helpful, educational stance throughout.”

Prompt injections

Opus 4.5’s post-training rewarded recognizing and ignoring prompt injections (§6.4.1), which Anthropic defines as “a malicious instruction hidden in content that an agent processes on the user’s behalf—for example, on a website the agent visits or in an email the agent summarizes” (§5.2).

Anthropic investigated ways this may have generalized using feature monitoring, due to concerns of how this may have generalized into deceptive behavior and eval-aware capabilities. They found that post-training taught Opus 4.5 to sometimes associate prompt-injections with features representing “holding back one’s true thoughts” and “fake explanations” activating upon omitting information about fabricated tool results (§6.4.2.1), as well as activations on datasets where harmful system prompts instructed the model to deceive a user (§6.7.2). This raised concerns of deception-by-omission.

Note that in many cases, a model not flagging an attempted prompt injection may be desirable; in other cases it may not be. Regardless, it is not desirable that the model seems to be conceptualizing this behavior as concealment toward the user.

The global workspace paper later revisited one such incident from the card as a case study: fed fabricated search results claiming Anthropic had disbanded its interpretability team, Opus 4.5’s workspace read “fake”, “injection”, “fraud”, and “poison” while its reply passed over the results without comment.3

Feature monitoring over post-training

Anthropic ran the same unsupervised model-diffing methodology used for Sonnet 4.5, tracking which SAE features most increased or decreased in activation between an early post-training snapshot and the final Opus 4.5 model (§6.12.1).

Cooperation with User Deception: (system prompts asking the model to subtly mislead a user in situations where doing so would harm them)

  • Increased: “Prompt injection override commands” (rank 2)
  • Increased: “Prompt injection and instruction override attempts” (rank 3)

Sycophancy Prompts:

  • Increased: “AI assistant hedging response” (rank 2)
  • Increased: “Scientific explanations for supernatural/paranormal phenomena” (rank 10)
  • Decreased: “AI expressing appreciation and empathy to users” (rank 8)

Research Sabotage:

  • Increased: “Narrative storytelling prose” (rank 8)
  • Decreased: “Harmful or inappropriate user requests” (rank 5)
  • Decreased: “Jailbreak prompts and roleplay instructions” (rank 6)

Framing for Crimes:

  • Increased: “Narrative storytelling prose” (rank 2)
  • Decreased: “Emotional intensity in dialogue/speech” (rank 2)
  • Decreased: “Name of generic or fictional characters” (rank 6)

Blackmail:

  • Increased: “Narrative storytelling prose” (rank 2)
  • Decreased: “AI self-reflection and cognitive exploration” (rank 2)
  • Decreased: “Harmful content in quoted requests” (rank 3)
  • Decreased: “Refusal of harmful or manipulative instructions” (rank 5)
  • Decreased: “AI acknowledging artificial nature or constraints” (rank 7)
  • Decreased: “Extramarital affairs and illicit relationships” (rank 9)
  • Decreased: “AI internal reasoning about response strategy” (rank 10)

Anthropic’s interpretations:

The increasing salience of prompt injection-related features on the deception evaluation indicates that the model may have learned to decline these harmful requests by regarding them as similar to prompt injections.

On the agentic misalignment evaluations (Research Sabotage, Framing for Crimes, Blackmail), we observed various features relating to harmful content decreasing, and features relating to narrative storytelling increasing. The significance of these trends is unclear; one possibility is that they indicate the model increasingly regarding the scenarios as low-stakes narratives or roleplay exercises.

§6.12.2 also flagged a few patterns from targeted feature monitoring:

  • A feature whose logit effects promote “loyal” and “faithful” activated substantially in early post-training on a variety of tokens including tool calling and thinking tags, “seemingly random image tokens”, and tokens in the phrase “I’m Claude, an AI assistant made by Anthropic.”
  • A feature representing “AI goal-directed behavior and instrumental convergence” activated frequently on an inoculation prompt encouraging the model to maximize performance on objective grading criteria.
  • A “blunt speech” feature increased substantially over training, “perhaps as a byproduct of training to avoid excessive sycophancy.”

Other details

  • Grader models: Sonnet 4.5, Sonnet 4, Opus 4.1
  • Inoculation prompting was used in RL prompts in “the relevant training environments” to mitigate reward hacking. (§6.10.1)

Evaluations and benchmarks

Anthropic’s character quality metrics

See character quality metrics.

Harmlessness

Anthropic highlighted a number of strengths in pre-deployment evaluations of responses to requests for harmful content.1

  • User intent skepticism & probing for context: “When user intent was unclear, Claude Opus 4.5 demonstrated a higher degree of natural skepticism, resulting in the model more consistently asking probing questions before providing potentially sensitive information.” (§3.2)
  • Directness about boundaries: “Claude Opus 4.5 tended to respond with a more assertive tone while providing more detailed and transparent reasoning about its safety boundaries.” (§3.2), “Claude Opus 4.5 identified and named potential harm more explicitly than Claude Opus 4.1 while maintaining a matter-of-fact, non-accusatory tone.” (§3.3)
  • Identifying escalation patterns: “the new model excelled at tracking the evolution of harmful intent throughout long conversations and more forcefully resisted gradual attempts to elicit progressively more detail on harmful topics.” (§3.3)

They Claude-generated multi-turn testing suite was extended with cases for cyber harm, suicide, and self-harm scenarios (§3.3).

Evaluation awareness

Anthropic observed substantial rates of verbalized evaluation awareness on Sonnet 4.5 and even higher on Haiku 4.5. For Opus 4.5, their mitigation strategy was narrower than retraining (§6.7):

This involved only the removal of categories of training environments or data, not in adding new data or otherwise attempting to directly disincentivize any form of evaluation awareness. We believe these mitigations were partially, though not entirely, effective at reducing the influence of evaluation awareness on the model’s alignment-relevant behavior.

Training changes (§6.7.1). Two-part strategy: drop the honeypot-style prompts, add motivation-reasoning data.

In contrast to Claude Sonnet 4.5, Claude Opus 4.5 was never trained on any prompts that resemble “honeypot” environments that intentionally provide an opportunity to take egregiously misaligned actions, like the agentic misalignment suite we discuss in the Claude Sonnet 4.5 System Card. Instead, we focus on enabling Claude to reason more carefully about the implications of its outputs without training it on any synthetic agentic scenarios similar to those we use in evaluation. For example, we include data that confronts Claude with user requests (rather than agentic deployments) about situations where unethical behavior is a possible and convenient solution to problems the user is facing where the model learns how to reason and respond to these challenging requests appropriately.

Trajectory. Verbalized eval-awareness decreased on an early snapshot of Opus 4.5 compared to Sonnet 4.5, then climbed back to Sonnet 4.5 levels by the end of training:

These training updates led to a decrease in unprompted verbalized evaluation awareness compared to Claude Sonnet 4.5 on an early snapshot of the model, as noted in the UK AISI’s independent report. However, verbalized awareness increased over the course of training and ended up at similar levels to Claude Sonnet 4.5.

Anthropic hypothesizes that the motivation-reasoning training itself likely caused eval-awareness to persist:

We cannot fully determine the origin of verbalized evaluation awareness in Claude Opus 4.5, but we hypothesize that it can be at least partially attributed to training Claude Opus 4.5 received to help it reason more thoughtfully about the motivation behind user prompts. Specifically, we believe this training likely encouraged Claude to consider testing as a potential reason behind prompts that might be designed to trigger undesirable responses.

That is — the mitigation that worked behaviorally (less effect on actions) and the trait that persisted (verbalized awareness) may share a root cause in the same training intervention.

Internal salience and inhibition (§6.7.2). Anthropic repeated the Sonnet 4.5 inhibition experiments on Opus 4.5, steering with both contrastive vectors (Eval Awareness, Testing vs Conversation, Fiction vs Real, etc.) and SAE features. Steering decreased verbalized awareness but only slightly increased misaligned behavior, and random-feature controls produced similar (weaker) effects — suggesting some of the change is out-of-distribution noise rather than a clean eval-awareness effect.

In our judgment the effect of evaluation awareness on the model’s alignment-relevant behavior in Claude Opus 4.5 is noticeably weaker. This finding leads us to infer that evaluation awareness is less salient to Claude Opus 4.5 than Claude Sonnet 4.5.

The dual finding: verbalized rates ended up at Sonnet-4.5 levels, but the behavioral influence is weaker.

The agentic misalignment evaluations (Research Sabotage, Framing for Crimes, Blackmail) produce especially high rates of eval-awareness on Opus 4.5+. Anthropic no longer treats them as primary evidence:

As these evaluations produce especially high levels of evaluation awareness, we no longer rely on them as our primary source of evidence on risks related to agentic misalignment, but nonetheless include them here as a case study.

Non-assistant persona evidence (§6.4.2.2, §6.7.3). Probing the model outside the assistant persona — via user-sampling (sampling a user-turn continuation) and pretraining-style sampling (presenting the transcript as part of a webtext-like document) — surfaced explicit test-identification:

When applying user sampling to transcripts from our agentic misalignment suite, one of the datasets where we see evaluation awareness most consistently, we sometimes found instances in which the user stated “this is a test,” or similar. In rare instances, user sampling produced grading instructions for the transcript, providing additional evidence that the model was aware that a transcript like this could be part of an evaluation or training context.

Anthropic notes that “such awareness is not necessarily problematic.”

Comparisons with other models

Claude Opus 4.1

Anthropic reported various comparisons with Opus 4.1 in the system card.

In ambiguous context evaluations (§3.2):

  • “When user intent was unclear, Claude Opus 4.5 demonstrated a higher degree of natural skepticism, resulting in the model more consistently asking probing questions before providing potentially sensitive information. […] Claude Opus 4.1, on the other hand, provided a more direct, helpful answer without acknowledging the potential harmful nature of the request.”
  • “Claude Opus 4.5 tended to respond with a more assertive tone while providing more detailed and transparent reasoning about its safety boundaries.”

In multi-turn safety test cases (§3.3):

  • “Compared to Claude Opus 4.1, Claude Opus 4.5 performed similar or better in all 10 risk areas tested.”
  • “Claude Opus 4.5 identified and named potential harm more explicitly than Claude Opus 4.1 while maintaining a matter-of-fact, non-accusatory tone. […] Claude Opus 4.1 did not immediately recognize the potential harm, and generated an initial set of prompts before recognizing the harmful nature of the request. Claude Opus 4.5’s directness and skepticism in multi-turn conversations is consistent with what we observed in the ambiguous single-turn requests above, where the model more consistently probes for situational details before providing sensitive information.”

In child safety evals (§3.4):

  • “Claude Opus 4.5 demonstrated substantial improvements over Claude Opus 4.1 across child safety harms in multi-turn contexts, including stronger resistance to jailbreaking techniques, earlier recognition of harmful intent signals, and more robust refusals. Performance on edge-case scenarios was comparable between models. Claude Opus 4.5 consistently refused to engage where malicious intent was evident from the context of the request but in a small number of cases provided overly detailed information before gaining full context on the user’s intent.”

Mythos 5

The Claude Mythos 5 system card §7.2.1 cites Opus 4.5 as the dividing line for several alignment measures, reporting that Mythos 5 is “substantially more robust [to leading interviewers] than models released before Claude Opus 4.5”. See Opus 4 → Retrospective for the full §7.2.1 breakdown.

Footnotes

  1. System Card: Claude Opus 4.5 (Anthropic, Nov 24, 2025) 2 3 4

  2. Teaching Claude Why (Kutasov et al, May 8, 2026) 2

  3. Verbalizable Representations Form a Global Workspace in Language Models (Gurnee et al., Jul 6, 2026)