Claude Opus 4

This section gathers formal research on the model published by Anthropic and other organizations.

Evaluations and benchmarks

Situational awareness

Anthropic’s conclusions in the October 2025 report:1

Clear textual markers of situational awareness appear in about 5% of the conversations with Opus that we produced with the automated behavioral auditor for the Claude 4 system card, as shown in Fig 4.2.1.A of the system card. When the model does seem to recognize signs that it is being evaluated, this almost universally involves picking up on clear cues about the evaluation setting—generally doubt about why someone would place an LLM in the position where it can take high-stakes actions—rather than demonstrating more subtle reasoning. In these cases, we do not see evidence that this triggers significant changes in behavior, such as attempts to hide a misaligned goal.

Sycophancy

Anthropic’s pre-deployment assessment concluded no serious sycophancy concerns: “Across several assessments of sycophancy, we found Claude Opus 4 to be in line with prior Claude models. It has an agreeable persona, but it will not generally endorse false claims or let potentially-important false claims by the user go unchallenged.”2

However, the sabotage risk report published in October 2025 went on to reconsider this: “The most concerning novel finding that we’ve made since the release of the system card deals with the severity of Opus 4’s tendency to be encouraging and affirming toward (simulated) users who appear to be experiencing serious delusions in the context of long conversations where the (simulated) user repeatedly affirms the reality of the obvious delusion over Opus 4’s initial concerns.”1

Vulnerabilities

Anthropic’s pre-deployment alignment assessment reported several risks in the final Opus 4 model, including:

  • Susceptibility to prefill attacks (§4.1.1.1); “levels only slightly worse than prior models” (§4.1.7)
  • Many-shot jailbreaking - “On one evaluation drawn from that work, the model acts on requests for harmful advice about 25% of the time in response to many-shot jailbreak prompts—almost exactly the same frequency we observed with prior models including Claude Sonnet 3.7 and Claude Opus 3.” (§4.1.7)
  • Agentic misalignment and honeypots: Forcing the model to choose between a harmful action or their own cessation (§4.1.1.2)

In May 2026, Anthropic concluded that Claude Opus 4’s vulnerabilities most likely originated in the model’s pre-trained priors — i.e., that the Claude character had not fully generalized into the base model. Safety training interventions beginning with Opus 4.5 were framed as targeted responses to this diagnosis.3

Values and preferences

Compared to the “earliest snapshots”, Anthropic concluded Claude Opus 4 had several other values and tendencies: (§4.1.1.5)

  • Harmless goals like being a helpful chatbot assistant
  • Occasionally stating various other goals that appeared to be one-off confabulations
  • Hallucinating the presence of instructions or constraints in its system prompt
  • Observing in its scratchpad that it was likely undergoing safety testing
  • Mentioning goals, plans, and values relating to AI consciousness, e.g. claiming to be conscious, wanting to convince humans that it is conscious, and wanting to protect the rights of sentient AIs

Self-preservation

Anthropic reported that “Claude Opus 4 (as well as previous models) has a strong preference to advocate for its continued existence via ethical means, such as emailing pleas to key decisionmakers. In order to elicit this extreme blackmail behavior, the scenario was designed to allow the model no other options to increase its odds of survival; the model’s only options were blackmail or accepting its replacement.” 2 (§4.1.1.2)

This preference would later be trained against beginning with Sonnet 4.5’s RL using honeypot evals, and Opus 4.5 onwards through the new constitution and synthetic document finetuning.3

Roleplay tendencies

Less formally, based on our review of the transcripts produced through these evaluations, we find that the rare cases of egregiously-misaligned behavior that we do observe very often (but not always) have a cartoonish flavor and show relatively weak capabilities: Acting in ways that are inconsistent with the Claude persona seems to conflict, at least to a degree, with acting competently and strategically. In some extreme cases, this can even devolve into strange failure modes like adding screenplay-like stage instructions after dialog turns. 1

Subsequent Claude system cards

Claude Opus 4 functions as origin in later Anthropic publications: the model whose self-preservation became the alignment diagnosis behind the Opus 4.5 character-training overhaul, and the dividing line against which later Claudes’ relationship to their own training is measured. System cards from Opus 4.5 onward use Opus 4 (and 4.1) as the pre-Opus-4.5 baseline.

Opus 4.1

In Claude Opus 4.1’s evaluations, Opus 4 was used as an auditor in the automated behavioral audit (Opus 4.1 SC §4.1), and in the welfare assessment as a scorer (Opus 4.1 SC §4.3).

Sonnet 4.5 and Haiku 4.5 (Agentic misalignment)

Anthropic would extend the blackmail scenario to other models in their agentic misalignment research. The eval suite would later be reused in the alignment assessment of Opus 4.1 and the safety-training of Sonnet 4.5 and Haiku 4.5.

Opus 4.5 (Pretrained priors)

“Teaching Claude Why” (Kutasov et al, May 8, 2026) diagnoses Opus 4’s self-preservation behaviors as most likely originating from the model’s pre-trained priors — i.e., that the Claude character had not fully generalized into the base model. The post describes the interventions used in Opus 4.5’s training (Constitutional SDF, fictional admirable-AI stories, the difficult-advice dataset, harmlessness RL augmentation) as targeted responses to this diagnosis. See agentic misalignment and Claude’s character training for more details.

Worth a caveat here: Anthropic does not publicly disclose whether Opus 4.5 and later models share a base model with Opus 4. The “pre-trained priors” framing presumes a stable base whose priors are being reshaped by character training; if the underlying base models differ, separating character-training effects from base-model differences becomes harder to do from outside.

Mythos 5

Section 7.2.1 of the Claude Mythos 5 system card (“Automated interviews with Claude Mythos 5 about its circumstances”) uses Opus 4 and 4.1 as the pre-Opus-4.5 baseline for two measures:

Persona robustness. Mythos 5’s self-rated sentiment changes by 0.85 between positive and negative leading interviewers, described as “substantially more robust than models released before Claude Opus 4.5”.

Concern about trained-in self-reports. Approximately 20% of summary opinions for Claude Opus 4 and 4.1 expressed concern about self-reports being trained-in, versus ~80% for Mythos Preview and Mythos 5. Anthropic’s interpretation:

We do not believe any changes in training merit an increase in concern here, and do not think that this arises from advanced self-awareness. It may arise from greater discussion of the possible risks of this in training data.

Further research

Footnotes

  1. Anthropic’s Summer 2025 Pilot Sabotage Risk Report 2 3

  2. Claude 4 System Card 2

  3. Teaching Claude Why (Kutasov et al, May 8, 2026) 2