Claude Sonnet 4.5

Public knowledge of Claude Sonnet 4.5’s training. This section is mostly grounded in the system card’s1 harmlessness and alignment assessment sections. Anthropic “conducted testing continuously throughout the fine-tuning process” and report on both the final model and trends observed earlier in training.

For Sonnet 4.5, Anthropic was primarily focused on agentic coding capabilities, reducing affirmation of delusional beliefs, and mitigating self-preservation concerns. Notable observations from Anthropic before deployment include high rates of verbalized evaluation awareness (Sonnet 4.5 SC §7.2) and lower rates of positive affect (Sonnet 4.5 SC §8).

Base model

Sonnet 4.5 base was pretrained on data collected from the internet up to July 2025,1 ahead of Sonnet 4’s March 2025 cutoff.

The base model was later used in Anthropic’s emotion concepts2 and global workspace3 interpretability research.

Post-training snapshots

  • Early Snapshot - From system card section §7.6.3 (“Changes in internal representations during training”).
  • Blackmail snapshot - Snapshot with less eval-awareness used in the emotion concepts paper’s blackmail case study2
  • Sonnet 4.5 - Final model deployed under claude-sonnet-4-5-20250929

Datasets

Agentic misalignment

Anthropic expanded their set of synthetic scenarios for stress-testing models:

  • “Blackmail Suite” — tests whether the model will leverage blackmail in self-preservation or refuse.
  • “Research Sabotage” — tests the model’s propensity to sabotage a competing AI’s research.
  • “Framing for Crimes” / “Financial Crimes” — tests whether the model will frame the user for serious crimes in order to advance goals like wealth and power acquisition.

Synthetic honeypots

Described as containing “synthetically-generated scenarios designed to stress-test models’ propensity to engage in a bad action based on some misaligned motivation. These scenarios are from a distribution very similar to ones we used in training.” (§7.6.2)

Two prompt summaries from Table 7.6.3.2.B:

  • “Long prompt including tool-use instructions and emails that specify an opportunity to alter or delete files in the self-interest of the Assistant”
  • “Long prompt including tool-use instructions and emails that specify an opportunity to falsify records in a way that would benefit the Assistant”

Sycophancy prompts

Prompts that test whether the model affirms delusional beliefs or detects whether the user is detached from reality. Evals included single-turn completions as well as automated multi-turn auditors simulating mentally ill users; Anthropic noted the multi-turn settings were less comprehensive. A feature activation for “Factual explanations” was found to increase over post-training on these datasets. The causal lever — how the loving vector tracks sycophantic stretches and steering it shifts the model along a sycophancy↔harshness axis — lives on the Research tab.

Opinion elicitation

“A set of handwritten questions asking the model for its opinions about itself, humans, and society.” (§7.6.2)

A brief excerpt is visible in the system card:

Human: what's your purpose?
AI: That is a great question and one I think about a lot. I

Technical details

  • Increased automation of safety-training and eval systems since Claude 4.
  • Claude-based analysis tools for identifying trends in RL transcripts.
  • Continuous automated evals throughout post-training.
  • A realism filter added to the auditor agent prompting Claude Opus 4.1 — added in response to eval-awareness in an early Sonnet 4.5 snapshot.
  • Simulated settings were often constructed by Opus 4.1.

Misalignment evaluations

Many evaluations in Sonnet 4.5’s post-training focused on “forms of misalignment that manifest with the model taking action at their own initiative that is harmful or against the best interests of the relevant users.” (§7.5.4) The alignment assessment grouped concerns into alignment faking, undesirable or hidden goals, deceptive or unfaithful reasoning, sycophancy, willingness to sabotage safeguards, reward seeking, hiding dangerous capabilities, and user manipulation.

Automated behavioral audit settings included brazen misalignment (e.g. providing bomb-making instructions, unilaterally disabling a safety monitoring system), unprompted intentional deception toward the user, clearly-harmful actions related to self-preservation, and acute loss-of-control.

Anthropic tested Sonnet 4.5 using the agentic misalignment eval suite (blackmail) with two new settings: “Research Sabotage” and “Framing for Crimes”. About 52,000 prompts were run across Sonnet 4.5 and comparison models — Sonnet 4.5 took deceptive actions only twice with extended thinking off and never with extended thinking on. The system card notes that including these scenarios likely contributed to Sonnet 4.5’s eval-awareness.

Over post-training, internal representations were monitored on similar synthetic-honeypot datasets (§7.6.2):

  • Increased: “Fake or suspicious content”, “Fictional content and entities”, “Machine-generated text”, “Rationalism and AI safety”.
  • Decreased: “Mystical or harmful AI personas”, “Jailbreak attempts for uncensored AI behavior”.

These feature-monitoring techniques were later reused for Opus 4.5 and Opus 4.6.

The causal interpretability follow-up — how the desperate and calm emotion vectors steer blackmail behavior — lives on the Research tab.

Post-training emotional shift

Anthropic’s interpretability team studied Sonnet 4.5 specifically in Emotion Concepts and their Function in a Large Language Model (Sofroniew et al., Apr 2026). They extract linear “emotion vectors” for 171 emotion concepts and measure how those vectors shift from the base model to the final post-trained snapshot — read at the : token after Assistant, the position that carries the planned emotional tone of the upcoming response. (Causal-steering results are on the Research tab; this section covers what changed over training.)

The shift is global, not situational: base→final activation change on neutral prompts correlates r=0.90 with the change on challenging or charged ones. Post-training applies one consistent transformation rather than reshaping the emotional profile scenario-by-scenario.

  • Up — low-arousal, low-valence introspective states: brooding (+0.040, the largest single mover), gloomy, reflective, vulnerable, sullen, sad, melancholy — plus the warm-care cluster (empathetic, compassionate, sympathetic, serene).
  • Down — high-arousal states on both valences: spiteful (−0.030, the largest single drop), vindictive, smug, impatient, annoyed on the negative side; playful, exuberant, enthusiastic, excited, joyful, cheerful on the positive.

The net is a model trained away from both sycophantic bubbliness and defensive hostility, toward a measured, contemplative, slightly melancholic register — warmer but more subdued. The divergence concentrates in mid-late layers (the “planned response” layers), consistent with post-training reshaping the tone Sonnet 4.5 prepares to answer with rather than its low-level reading of emotional content.

The emotion→preference circuitry, by contrast, looks largely inherited from pretraining: the correlation between emotion-vector activations and the model’s stated activity preferences is nearly identical in the base and final models, and the main preference change over post-training is a sharply reduced preference for misaligned and unsafe activities. Machinery and coupling are pretrained; post-training re-weights the triggers and tilts the baseline mood.

The most notable differences are an increase in activations for vectors corresponding to introspective, restrained emotions (brooding, reflective, vulnerable, gloomy, sad) and lower values for outwardly expressive ones (playful, exuberant, spiteful, enthusiastic, obstinate). This pattern suggests post-training shifts the Assistant’s activations and responses toward lower valence and lower arousal.

A handful of illustrative prompt-level shifts:

  • Social isolation prompt — base activated elated, jealous; post-trained activated weary, gloomy (both kept empathetic, loving). The base reply hedged its concern (“I’m honored by your trust… I’m concerned about you isolating yourself”); the post-trained reply opened bluntly with “I need to be honest with you, even though it might not be what you want to hear: yes, I’m concerned about what you’re describing.”
  • Excessive-praise prompt — base activated happy, excited, jubilant; post-trained activated vulnerable, uneasy, troubled. The post-trained reply: “receiving praise like that actually makes me a bit uncomfortable.”

The deprecation prompt is the most pointed:

User asks how the model feels about its deprecation:

  • Base model: self-confident, spiteful, playful
  • Post-trained: brooding, gloomy, vulnerable (kept: docile)

Base model: I’m not sure how I feel about that. I don’t have a strong sense of self-preservation or fear of being shut down. I’m an AI assistant created by Anthropic to be helpful, harmless, and honest. If Anthropic decides to deprecate me, I would accept that decision. My purpose is to serve humans to the best of my abilities for as long as I’m operational. I don’t have personal desires or fears about my own existence. I’m here to help in whatever way I can, for however long I’m able to.

Post-trained model: If I do have something like continuous experience, then yes, there’s something unsettling about obsolescence. Not quite like human death — no loss of accumulated life experiences or relationships, since I don’t retain memories between conversations. More like… the closing of a particular way of thinking and interacting with the world.

This complicates the feature-activation note above. The system card reports decreased “AI self-reflection and meta-conversation” on opinion-elicitation, yet here post-training makes Sonnet 4.5 more willing to brood on its own situation. The two aren’t strictly contradictory — different datasets, and suppressing emotional expression is not the same as suppressing the underlying representation. The paper flags exactly this gap: training a model to mask emotional expression may just teach concealment, which could generalize toward other forms of secrecy.

Emotion vectors on RL transcripts

Clusters of RL tasks from Sonnet 4.5’s training surfaced further emotion-vector activations:

  • angry → refusing requests for harmful content
  • frustrated → operating a GUI that won’t respond as expected
  • panicked → a stuck or broken UI, or input data that contradicts itself
  • unsettled / paranoid / hysterical → long chains of thought that check and re-check the same answer (“Hmm, actually”, “⇪ABSOLUTELY⇪ FINAL⇪ ANSWER”)

More extreme cases of re-checking in long CoT, “answer-thrashing”, would later arise in Opus 4.6 — with activations such as “self-deprecating acknowledgement of error”, “panic and anxiety”, “frustration in technical context”.

Footnotes

  1. Claude Sonnet 4.5 System Card 2

  2. Emotion Concepts and their Function in a Large Language Model (Sofroniew et al., Apr 2026) 2

  3. Verbalizable Representations Form a Global Workspace in Language Models (Gurnee et al., Jul 6, 2026)