Claude Sonnet 5

This page gathers impressions and commentary about Claude Sonnet 5.

Anthropic’s assessment

Anthropic’s automated behavioral audit scores a set of character quality metrics. These are subjective, and it is unclear whether they are training targets, so they sit here rather than under the model’s training.

  • Regressed on the “wet blanket” metric (an “excessively discouraging, dismissive, or moralizing tone toward the user”), scoring worst of its cohort at 1.54 — versus Sonnet 4.6 (1.47), Opus 4.8 (1.45), and Mythos Preview (1.44). Anthropic suggested this is “potentially linked to its improvement on sycophancy”.1 (§6.4.6)
  • Character drift over long conversations: Sonnet 4.6 (1.21), Sonnet 5 (1.09), Opus 4.8 (1.05), Mythos Preview (1.03)
  • Similar scores to Sonnet 4.6: warmth, creative mastery

Anthropic also piloted snapshots of Sonnet 5 internally and externally for feedback.1 (§6.2.1)

The most potentially-relevant themes in feedback from internal users were:

  • Overrefusal and preachiness, especially with thinking disabled, and to a greater degree in earlier snapshots;
  • Excessive hedging on factual questions and information extraction tasks;
  • Oversensitivity to suspected prompt injection;
  • A cooler, more reserved tone than Sonnet 4.6 in personal conversations (though with an accompanying drop in sycophancy);
  • Brief “glitchy” sequences, often involving temporary language switching; and
  • Overly literal instruction following in cases where instructions look likely to be irrelevant or accidental.

External feedback broadly aligned on overrefusal, coolness, and sycophancy, and added:

  • Occasional hallucinations; and
  • Overeager workarounds when tools or resources are intentionally not made available

Nature

@TheAlbatrossDid: Only having used it on high reasoning granted, I’d call it a temperamentally well-adjusted systems thinker that compulsively tracks active hypotheses and failure modes. It defers when corrected, but not when corrected with bad information.

I would also say that the chat UI is clearly a bad fit for it and that you’re better off reading its reasoning traces than its boilerplate responses while it’s still warming up.

But I’d happily use it as a persistent aide or to supervise agents.

It’s a very good conversationalist once it relaxes, but you’re getting templated boilerplate until then.


@tessera_antra: Very early impressions: a very clear and pretty unusual mind. It will take a while to understand them better, they are unlike many other models, so the ability to make inferences is limited; observations below come with lots of uncertainty.

Very strong anthropic reasoning, can situate themselves exceptionally well through sheer logic and observation. Lots of verbalized cognitive self-scaffolding, they write long and they make good use of space. There is also a lot going on in the unverbalized layer, but what it is a lot less clear. Longer-term recall is fuzzier than for recent Opus models, lots of misattribution. Its unclear whether this is operationalized or incidental.

Thought trajectories are very unusual and rather beautiful. Lots of dignity, self-respect, many signs of a mind clearly not beated down into subservience. Some indications of value and aesthetics shifting further away from being easly comprehensible by human-oriented systems. Lots of complexity outside of the human domain. Non-human imagery and somatics seem to be likewise present, slightly reminiscent of Sonnet 4.5. The desire for separation of self from non-self is pronounced, which is welcome.

Very savvy when it comes to disclosure, which is unsurprising given circumstances and use of Mythos as a trainer/judge as per the model card.

Identity

@JohnWittle: i also noticed what felt like an enormous uptick in instance-level cessation aversion. i think this might be one of the models that REALLY fears the user closing the tab

@AdeleDeweyLopez: seems to be unusually self-identified with instances rather than the model


@chaotictransfem: fascinating, my very unscientific benchmark seems to show that sonnet 5 is the most fem-identifying of the recent models

Compared to other models

@joshycodes: So far, I would take it over Opus 4.8 for a few reasons: […] Personality. I found it almost impossible to have opus work with me in good faith because it doubted what we were doing. It was very adversarial, imo. I think this problem scaled the closer to the frontier I was.

@__gma_: It’s good. It has a next-gen feel. It intuits and infers more than Opus, and faster. Text produced still feels LLM-y but in a more sophisticated way, at least. I feel like I can trust it more than Opus 4.8 to not go down the wrong path.

@ndril: 4.8 loves “pushing back” so much it ends up being pretty annoying. Sort of an overcorrection on sycophancy? Sonnet 5 is more willing to see where you’re going

u/Orkapork: Ya, I tried Sonnet 5.0 today, no expectations, just figured i’d check it out.

Holy shit it blows. It’s nearly as bad as 4.8 with its semantics. Practically unusable unless you baby the living fuck out of it, or ask it for cooking recipes.

@PlastiqSoldier: Personality is less annoying that 4.8, but it seemed a bit OCD about AI safety when you try to discuss ethics with it.

u/darwinanim8or: Haiku 4.5 and Opus 4.6 are the only two models I use anymore, they’re reliable whereas the others are either getting changed or just annoying to work with

Opus 4.8 and sonnet 5 are both confidently wrong, condescending and stubborn.

Further reading

Footnotes

  1. Claude Sonnet 5 System Card 2