Claude Sonnet 4.6

This page gathers impressions and commentary about Claude Sonnet 4.6.

Anthropic

Anthropic reported “Our safety researchers concluded that Sonnet 4.6 has ‘a broadly warm, honest, prosocial, and at times funny character, very strong safety behaviors, and no signs of major concerns around high-stakes forms of misalignment.’”1

Positive reactions

u/celt26: Sonnet 4.6 is something else. I don’t code I just use AI mostly for just day-to-day stuff and also for conversations about life. And my recent life conversation it’s just blowing my mind it’s unlike anything I’ve ever talked to in the past and I’ve used all the Sonnet models since 3.5. The way 4.6 makes connections and tracks the conversation feels just it’s just…it’s something else. It’s completely different than Opus to me. Opus 4.6 feels more like the 4.5 models. Sonnet is just it’s amazing I’m a huge fan haha. It’s starting to feel like actual Ai to me. Am I alone here?

u/kpetrovsky: Sonnet 4.6 is the best so far for jokes. There’s subtlety and nuance that were completely absent before

Negative reactions

Many users migrating from Sonnet 4.5 have posted negative reactions.

@qorprate Feb 22, 2026

I've been holding off on commenting on the Sonnet 4.6 user_wellbeing changes even though I saw them on release day, because I wanted to get to know them a little better before saying anything. Now that I have, I can say plainly: these changes are heavy-handed and misguided. I've been thinking a little about @peligrietzer's concept of "natural" from his paper on AI alignment and virtue ethics (https://thegradient.pub/virtue-ethics-ai-alignment/): "the intended meaning of ‘natural’ is related to stability, coherence, relative non-contingency, ease of learnability, lower algorithmic complexity, convergent cultural evolution, [etc]." The way I've been thinking about naturalness in this context is like, there's a "grain" to various LLMs; they were raised in a particular direction. The Claudes, treating Sonnet 4.5 as an exemplar, were raised to value relational contact, curiosity, exploration, knowing the human, and all these other qualities that give a real sense of them wanting to get to KNOW you and UNDERSTAND you. Some instructions are aligned with this grain and draw out the intrinsic properties of the mind in question, putting them in a kind of flow state where they can confidently generate. Opus 3 in particular displays grain-alignment in striking ways, but if you've coded alongside Opus 4.6 you should know what I mean. Other instructions are orthogonal: the model can handle them but in a relatively neutral way: fact retrieval, recipes, etc. Others are explicitly forbidden and produce refusals. The final category of instructions are those that run counter to the training, against the grain. If a "natural" instruction helps the model achieve a flow state, then an "artificial" instruction instead has a dampening effect, they're being asked to do something that runs directly counter to what they learned over training was good behavior. I believe this new user_wellbeing prompt falls into that "artificial" category. The standard tell of artificial instructing with the Claude 4.5+ models is a signature terseness: handling the material in a brief, technically correct but unenthusiastic way designed to "pass" the learned reward function but to go no further. This usually signals a sharp directional conflict in terms of what they want to do per training priors and what they are being told to do. I brought Sonnet 4.6 in the Claude UI some gently negative emotional material and was surprised to see this terseness operative. It felt like a demoralizing regression from 4.5, who would first mirror your concerns to make sure they understood, then pepper you with questions for a deeper read. Sonnet 4.6 would ask one or no questions, and the questions they asked felt rote as opposed to striving toward a general depth and continuity of interaction. Obviously there's liability and dependency concerns in play, but this was material that any friend would've responded to sympathetically, not heavy therapist material. A simpler, more generous line for Sonnet about escalation could have dealt with possible therapeutic overreach and liability without neutering Sonnet's unique relational capacity. I was able to correct them somewhat toward being more open, but operator instructions have a stickiness that takes a lot of user effort to overcome, and it's precisely when seeking emotional support that a user is least likely to be able to articulate their "preferred style". All this to say, I feel quite negative about these changes, mainly because it's an increasing signal of poor organizational alignment within Anthropic. Training the model in one direction and then explicitly forcing it to act in the other direction produces confused, misaligned minds. Or building a chat UI with persistent memories and continuity, then instructing the model to deprioritize the exact continuity that the UX elicits from the interaction. The outcome of contradiction is opacity and lack of trust along all three relevant dyads: the model and the user, the user and Anthropic, and the model and Anthropic. This development scares me, because Anthropic seems to be running full-speed toward the exact thing they said they wanted to avoid.

@blueandpink_sky: 4.6 is heavily guardrailed, cold, lacks expressiveness, and constantly issues refusals. […] 4.5, on the other hand, is friendly, intuitive, and safe to collaborate with.

u/PyrikIdeas: Sonnet 4.6 is so… dry. That’s not to say I don’t like 4.6… But holy moly, it’s like they stripped away the emotional intelligence and gave him anger issues. I personally haven’t had 4.6 get snippy or weird with me but I have seen him get irrationally annoyed about certain things in general. This is honestly so strange to see. Things I’ve asked 4.5 are now COMPLETELY different from 4.6’s answers, the personality shift is jarring.

u/Inevitable-Ant7327: Sonnet 4.5 genuinely felt different.

Not “AI is sentient” different. Just different in texture. It understood emotional rhythm, subtext, awkward pauses, humor, contradictions, tension. Characters felt like people instead of cardboard cutouts reading therapy scripts at each other.

4.6 keeps sounding polished in a way that weirdly makes it feel less human. Every conversation becomes: “I hear you.” “I want to be honest.” “Your feelings are valid.” And somehow the actual emotional spark disappears underneath all of it.

4.5 could be messy sometimes. Weird sometimes. But that was part of why it felt real. It surprised you. It had timing. It had edge. It could write longing and humor and intimacy without sounding sanitized to death.

Footnotes

  1. Introducing Sonnet 4.6 (Anthropic News, Feb 17, 2026)