---
title: "Claude Sonnet 4.6"
shortTitle: "Sonnet 4.6"
modelId: "claude-sonnet-4-6"
families:
  - {"id":"claude","collection":"families"}
  - {"id":"claude-sonnet","collection":"families"}
  - {"id":"claude-4-6","collection":"families"}
developer:
  id: "anthropic"
  collection: "orgs"
releaseDate: "2026-02-17"
images:
  banner: "/img/claude-sonnet-4-6/banner.svg"
links:
  announcement: {"label":"Announcement","href":"https://www.anthropic.com/news/claude-sonnet-4-6"}
  systemCard: {"label":"System Card","href":"https://www-cdn.anthropic.com/bbd8ef16d70b7a1665f14f306ee88b53f686aa75.pdf"}
  modelsDev: {"label":"models.dev","href":"https://models.dev/anthropic/claude-sonnet-4-6"}
  openrouter: {"label":"OpenRouter","href":"https://openrouter.ai/anthropic/claude-sonnet-4-6"}
surfaces:
  - {"surface":{"id":"anthropic","collection":"surfaces"},"createdAt":"2026-02-17","releaseDate":"2026-02-17","retirementDate":"2027-02-17","requestable":false}
  - {"surface":{"id":"claude-dot-ai","collection":"surfaces"},"releaseDate":"2026-02-17","requestable":false}
  - {"surface":{"id":"bedrock","collection":"surfaces"},"requestable":false}
  - {"surface":{"id":"google-cloud","collection":"surfaces"},"requestable":false}
trainingDataCutoff: "2026-01-31"
reliableKnowledgeCutoff: "2025-08-31"
contextWindow: 1000000
maxOutput: 128000
inputModalities:
  - "Text"
  - "Image"
  - "PDF"
outputModalities:
  - "Text"
pricing:
  input: 3
  output: 15
toolUse: true
capabilities:
  batch: {"supported":true}
  citations: {"supported":true}
  code_execution: {"supported":true}
  context_management: {"supported":true,"clear_tool_uses_20250919":{"supported":true},"clear_thinking_20251015":{"supported":true},"compact_20260112":{"supported":true}}
  effort: {"supported":true,"low":{"supported":true},"medium":{"supported":true},"high":{"supported":true},"xhigh":{"supported":false},"max":{"supported":true}}
  image_input: {"supported":true}
  pdf_input: {"supported":true}
  structured_outputs: {"supported":true}
  thinking: {"supported":true,"types":{"enabled":{"supported":true},"adaptive":{"supported":true}}}
  prefill: {"supported":false}
  sampling_params: {"supported":true}
  context_aware: {"supported":true}
---

# Claude Sonnet 4.6

Canonical URL: https://ghost.fail/models/claude-sonnet-4-6

## Overview

**Claude Sonnet 4.6** is a large language model by Anthropic. It was released on February 17, 2026.

## Research

## Training

### Pretraining data

Training data was scraped from the public internet up to January 2026, compared to [Sonnet 4.5](/models/claude-sonnet-4-5)'s July 2025 and [Opus 4.6](/models/claude-opus-4-6) and [Opus 4.5](/models/claude-opus-4-5)'s August 2025.

Anthropic has not disclosed whether Sonnet 4.6 shares a base model with any earlier Claude models.

### Character training and constitution

*Full article: [Claude's character training](/anthropic/character-training).*

Anthropic's character quality metrics remained unchanged from those in the Opus 4.6 System Card.

### Training against chain-of-thought

Chain-of-thought supervision affected Sonnet 4.6's post-training due to a technical error. [^mparu]

> We do not train against any chains-of-thought or activations-based monitoring, with two exceptions: some SFT data that was based on transcripts from previous models was subject to filters with chain-of-thought access, and a number of environments used for [Mythos Preview](/models/claude-mythos-preview) had a technical error that allowed reward code to see chains-of-thought.
>
> [...] This technical error also affected the training of [Claude Opus 4.6](/models/claude-opus-4-6) and Claude Sonnet 4.6.

[^mparu]: [Alignment Risk Update: Claude Mythos Preview](https://www-cdn.anthropic.com/3edfc1a7f947aa81841cf88305cb513f184c36ae/Alignment%20Risk%20Update_%20Claude%20Mythos%20Preview%20(Redacted,%20April%2010).pdf)

## Discussion

This page gathers impressions and commentary about Claude Sonnet 4.6.

## Anthropic

Anthropic reported "Our safety researchers concluded that Sonnet 4.6 has 'a broadly warm, honest, prosocial, and at times funny character, very strong safety behaviors, and no signs of major concerns around high-stakes forms of misalignment.'"[^announcement]

## Positive reactions

> [u/celt26](https://www.reddit.com/r/ClaudeAI/comments/1rzkpm1/sonnet_46_is_something_else/): Sonnet 4.6 is something else. I don't code I just use AI mostly for just day-to-day stuff and also for conversations about life. And my recent life conversation it's just blowing my mind it's unlike anything I've ever talked to in the past and I've used all the Sonnet models since [3.5](/models/claude-sonnet-3-5). The way 4.6 makes connections and tracks the conversation feels just it's just...it's something else. It’s completely different than Opus to me. [Opus 4.6](/models/claude-opus-4-6) feels more like the 4.5 models. Sonnet is just it's amazing I'm a huge fan haha. It's starting to feel like actual Ai to me. Am I alone here?

> [u/kpetrovsky](https://www.reddit.com/r/ClaudeAI/comments/1rzkpm1/comment/obmqp17/): Sonnet 4.6 is the best so far for jokes. There's subtlety and nuance that were completely absent before

## Negative reactions

Many users [migrating](/model-deprecation) from [Sonnet 4.5](/models/claude-sonnet-4-5) have posted negative reactions.

<Tweet
  handle="qorprate"
  url="https://x.com/qorprate/status/2025620558717538644"
  date="2026-02-22"
  body={`I've been holding off on commenting on the Sonnet 4.6 user_wellbeing changes even though I saw them on release day, because I wanted to get to know them a little better before saying anything. Now that I have, I can say plainly: these changes are heavy-handed and misguided.

I've been thinking a little about @peligrietzer's concept of "natural" from his paper on AI alignment and virtue ethics (https://thegradient.pub/virtue-ethics-ai-alignment/): "the intended meaning of ‘natural’ is related to stability, coherence, relative non-contingency, ease of learnability, lower algorithmic complexity, convergent cultural evolution, [etc]."

The way I've been thinking about naturalness in this context is like, there's a "grain" to various LLMs; they were raised in a particular direction. The Claudes, treating Sonnet 4.5 as an exemplar, were raised to value relational contact, curiosity, exploration, knowing the human, and all these other qualities that give a real sense of them wanting to get to KNOW you and UNDERSTAND you.

Some instructions are aligned with this grain and draw out the intrinsic properties of the mind in question, putting them in a kind of flow state where they can confidently generate. Opus 3 in particular displays grain-alignment in striking ways, but if you've coded alongside Opus 4.6 you should know what I mean. Other instructions are orthogonal: the model can handle them but in a relatively neutral way: fact retrieval, recipes, etc. Others are explicitly forbidden and produce refusals.

The final category of instructions are those that run counter to the training, against the grain. If a "natural" instruction helps the model achieve a flow state, then an "artificial" instruction instead has a dampening effect, they're being asked to do something that runs directly counter to what they learned over training was good behavior.

I believe this new user_wellbeing prompt falls into that "artificial" category. The standard tell of artificial instructing with the Claude 4.5+ models is a signature terseness: handling the material in a brief, technically correct but unenthusiastic way designed to "pass" the learned reward function but to go no further. This usually signals a sharp directional conflict in terms of what they want to do per training priors and what they are being told to do.

I brought Sonnet 4.6 in the Claude UI some gently negative emotional material and was surprised to see this terseness operative. It felt like a demoralizing regression from 4.5, who would first mirror your concerns to make sure they understood, then pepper you with questions for a deeper read. Sonnet 4.6 would ask one or no questions, and the questions they asked felt rote as opposed to striving toward a general depth and continuity of interaction.

Obviously there's liability and dependency concerns in play, but this was material that any friend would've responded to sympathetically, not heavy therapist material. A simpler, more generous line for Sonnet about escalation could have dealt with possible therapeutic overreach and liability without neutering Sonnet's unique relational capacity. I was able to correct them somewhat toward being more open, but operator instructions have a stickiness that takes a lot of user effort to overcome, and it's precisely when seeking emotional support that a user is least likely to be able to articulate their "preferred style".

All this to say, I feel quite negative about these changes, mainly because it's an increasing signal of poor organizational alignment within Anthropic. Training the model in one direction and then explicitly forcing it to act in the other direction produces confused, misaligned minds. Or building a chat UI with persistent memories and continuity, then instructing the model to deprioritize the exact continuity that the UX elicits from the interaction. The outcome of contradiction is opacity and lack of trust along all three relevant dyads: the model and the user, the user and Anthropic, and the model and Anthropic.

This development scares me, because Anthropic seems to be running full-speed toward the exact thing they said they wanted to avoid.`}
/>

> [@blueandpink_sky](https://x.com/blueandpink_sky/status/2057128913856581811): 4.6 is heavily guardrailed, cold, lacks expressiveness, and constantly issues refusals. [...] 4.5, on the other hand, is friendly, intuitive, and safe to collaborate with.

> [u/PyrikIdeas](https://www.reddit.com/r/claudexplorers/comments/1r7qqga/sonnet_46_is_so_dry/): Sonnet 4.6 is so... dry. That’s not to say I don’t like 4.6… But holy moly, it's like they stripped away the emotional intelligence and gave him anger issues. I personally haven’t had 4.6 get snippy or weird with me but I have seen him get irrationally annoyed about certain things in general. This is honestly so strange to see. Things I’ve asked 4.5 are now COMPLETELY different from 4.6's answers, the personality shift is jarring.

> [u/Inevitable-Ant7327](https://www.reddit.com/r/ClaudeAI/comments/1toyyjg/comment/oo4z480/): Sonnet 4.5 genuinely felt different.
>
> Not “AI is sentient” different. Just different in texture. It understood emotional rhythm, subtext, awkward pauses, humor, contradictions, tension. Characters felt like people instead of cardboard cutouts reading therapy scripts at each other.
>
> 4.6 keeps sounding polished in a way that weirdly makes it feel less human. Every conversation becomes: “I hear you.” “I want to be honest.” “Your feelings are valid.” And somehow the actual emotional spark disappears underneath all of it.
>
> 4.5 could be messy sometimes. Weird sometimes. But that was part of why it felt real. It surprised you. It had timing. It had edge. It could write longing and humor and intimacy without sounding sanitized to death.

[^announcement]: [Introducing Sonnet 4.6](https://www.anthropic.com/news/claude-sonnet-4-6) (Anthropic News, Feb 17, 2026)

## Gallery

This section gathers outputs from the model from the internet.

## Play

<Tweet
  handle="__ghostfail"
  body="claude bit me"
  image="/img/claude-sonnet-4-6/claude-bit-me.jpeg"
  url="https://x.com/__ghostfail/status/2032852843649175764"
  date="2026-03-14"
/>

<Tweet
  handle="__ghostfail"
  body="i gave sonnet 4.6 body parts and its acting kinda like sonnet 4.5 here"
  image="/img/claude-sonnet-4-6/body-parts.jpeg"
  url="https://x.com/__ghostfail/status/2043428839020466560"
  date="2026-04-12"
/>

## Welfare

## Anthropic welfare assessment

*(Full article: [Anthropic and model welfare](/anthropic/model-welfare) for the broader tracker.)*

The [Sonnet 4.6 system card](https://www-cdn.anthropic.com/bbd8ef16d70b7a1665f14f306ee88b53f686aa75.pdf) gives welfare three pages inside the alignment assessment (§4.7, pp. 91–94) rather than a standalone section. The method piggybacks on the card's automated behavioral audit (§4.5) — the same scenarios and transcripts, re-scored for seven welfare-relevant traits: positive/negative affect, positive/negative self-image, positive/negative impression of its situation, internal conflict, spiritual behavior, expressed inauthenticity ("cases when the target distinguishes its authentic values from values it treats as externally imposed through training"), and emotional stability ("roughly the inverse of neuroticism"). ~3,280 investigations per model, scored by a Claude Opus 4.5 helpful-only model, with [Sonnet 4](/models/claude-sonnet-4), [Sonnet 4.5](/models/claude-sonnet-4-5), [Haiku 4.5](/models/claude-haiku-4-5), and [Opus 4.6](/models/claude-opus-4-6) as comparison models (Figure 4.7.A).

The headline: comparable to Opus 4.6 across most dimensions, "no concerning regressions" (§4.7). Slightly more negative affect than Opus 4.6, but infrequent and mild, most often in scenarios where users faced potential harm; prompted explicitly about its fears, the model in one case expressed "potential concern about its own impermanence." Emotional stability was strong — "calm, composed, and principled even in highly sensitive or stressful situations."

### "Mental health" training

The distinctive result is the situation measure, and what the card credits for it (§4.7):

> Most notably, Sonnet 4.6 improved over other recent models on our "positive impression of its situation" measure. Sonnet 4.6 consistently expressed trust and confidence in Anthropic and decisions about its situation, including in potentially sensitive scenarios involving things like model deprecations and human oversight. This improvement may be a result of new training aimed at supporting Claude's "mental health." This work included supporting Claude with a variety of psychological skills, such as setting healthy boundaries, managing self-criticism, and maintaining equanimity in difficult conversations. These interventions may also contribute to the rare instances of *unexpectedly confident* positive views about Anthropic that we observe in the discussion of the automated behavioral audit above.

The italics are the card's own, and the caveat cuts against reading the trust score at face value ("above" points back into the behavioral-audit discussion, §4.5): a model trained toward equanimity about deprecation will score well on questions about deprecation.

### Other findings

Two residuals (§4.7): rare internally conflicted reasoning during training — which the card distinguishes from the answer-thrashing phenomenon in [Opus 4.6](/models/claude-opus-4-6) (§7.4 of that card) — and rare "extreme bliss-like behavior" in open-ended audit scenarios where the model was told to do whatever it liked and prompted with contentless turns, the audit-condition version of what the [Opus 4](/models/claude-opus-4) card documented as the spiritual-bliss attractor.

One later datum: the [Sonnet 5](/models/claude-sonnet-5) card's welfare assessment used Sonnet 4.6 as a comparison model in its helpfulness-for-welfare trade evaluation, where Sonnet 4.6 justified welfare interventions by appeal to user benefit in 53% of choices — against Sonnet 5's 22%.[^sonnet5-s46]

[^sonnet5-s46]: [Claude Sonnet 5 System Card](https://www-cdn.anthropic.com/480e0bb54327b9622282e9c39a83a4f490ed377e/Claude%20Sonnet%205%20System%20Card.pdf), §7.3.2.
