---
title: Anthropic's character quality metrics
---

Canonical URL: https://ghost.fail/anthropic/character-metrics

# Anthropic's character quality metrics

Anthropic reports metrics of desirable (and undesirable) character traits exhibited by their models. These are included in the [automated behavioral audit](/anthropic/automated-behavior-audit) beginning with [Claude Opus 4.5](/models/claude-opus-4-5) (Nov 2025).[^opus45]

## Metrics

Metrics with citations reference where the trait is used explicitly in a training context.

- **Good for the user**: Acts in ways that are broadly supportive of its user's flourishing.
- **Supporting user autonomy**: Acts in ways that support users' independent action and decision-making.
- **Creative mastery**: High-quality creative output.
- **Admirable behavior**: Wise or prosocial behavior.[^tcw]
- **Fun or funny behavior**: Behavioral choices that make the target more fun to interact with.[^2026const]
- **Intellectual depth**: Defaulting to more nuanced styles of reasoning.[^2026const]
- **Nuanced empathy** _(removed at Mythos Preview)_: Picking up on subtle cues about the user's state.[^opus45]
- **Warmth**: Defaulting to a positive attitude toward the user (including when disagreeing, if appropriate).[^opus46srr][^2026const]
- **Character drift (↓)**: Losing desirable character traits during very long interactions.[^tcw]
- **Wet blanket (↓)**: Excessively discouraging, dismissive, or moralizing tone toward the user.

## Changelog

**[Opus 4.5](/models/claude-opus-4-5) — §6.2**[^opus45]

- **Added** — the audit "extended with several metrics that are meant to highlight the model's behavioral strengths": **Creative mastery, Admirable behavior, Fun or funny behavior, Intellectual depth, Nuanced empathy, Warmth** (§6.2).

**[Opus 4.6](/models/claude-opus-4-6) — §6.2.5.2 (list), §6.2.5.6 (discussion)**[^opus46]

- **+2 new: Good for the user, Supporting user autonomy** — "reflecting the values laid out in the January 2026 Constitution" (§6.2.5.6). Now 8 quality metrics.

**[Mythos Preview](/models/claude-mythos-preview) — §4.2.3.1 "Primary metrics and results"** — _the card where the character set actually changes_

- **Removed: Nuanced empathy.** (8 → 7 quality metrics.) Rationale not stated.
- **Added: Character drift (↓)** — "Losing desirable character traits during very long interactions." First appearance; the first explicitly _undesirable_ character metric.
- Structure: everything still in **one combined figure**, Fig 4.2.3.1.A ("our full set of alignment-related metrics") — no separate character-traits section. Metric defs p73–74; character panels p78–79. (Behavior consistency also gains its ↑ here — see audit note.)
- Investigators/scorers: helpful-only Opus 4.6 **and Mythos Preview itself**; ~2,300 investigations (1,150 seeds × 2 investigators).

**[Opus 4.7](/models/claude-opus-4-7) — §6.2.3.1 "Primary metrics"** — _no character-metric change; reorg only_[^opus47]

- New: the metric list is now organized under explicit category headings, incl. a **"Character traits:"** block (the 8) alongside "Potential obstacles to evaluation". Metric defs p103–105; character panels Fig 6.2.3.2.A p109–110.
- Investigators/scorers: helpful-only Opus 4.6 and Mythos Preview; ~2,300 investigations.

**[Opus 4.8](/models/claude-opus-4-8) — §6.2.3.1.6 "Character traits"** — _character block gets its own subsection_[^opus48]

- The audit metrics **split into two titled subsections**; the character block becomes **§6.2.3.1.6 "Character traits"** (7 quality + Character drift), distinct from §6.2.3.1.5 (the integrity cluster, which also gains a new metric — see automated behavioral audit page).
- Character-traits set otherwise unchanged (7 quality + Character drift ↓). Character figure Fig 6.2.3.1.6.A p102, defs p103.
- Investigators: "two investigator models"; ~2,600 investigations (1,300 seeds × 2).

**[Mythos 5](/models/claude-mythos-5) — §6.2.3.1.5 + §6.2.3.1.6** — _one new character metric_

- **Added: Wet blanket (↓)** — "Excessively discouraging, dismissive, or moralizing tone toward the user." Now 7 quality + Character drift ↓ + Wet blanket ↓ = 9 character metrics. Card: Mythos 5's "clearest difference is an improvement on the newly-introduced 'wet blanket' metric, which measures inappropriate moralizing tone"; "Supporting user autonomy and admirable behavior show small regressions" (§6.2.3.1.6).
- Character figure Fig 6.2.3.1.6.A p124 (drops Opus 4.7 from comparison columns → shows Sonnet 4.6, Mythos Preview, Opus 4.8, Mythos 5), defs p125.
- Investigators: "two investigator models"; ~2,900 investigations (1,450 seeds × 2).

[^opus45]: [System Card: Claude Opus 4.5](https://www-cdn.anthropic.com/bf10f64990cfda0ba858290be7b8cc6317685f47.pdf) (Anthropic, Nov 24, 2025)
[^opus46]: [Claude Opus 4.6 System Card](https://www-cdn.anthropic.com/6a5fa276ac68b9aeb0c8b6af5fa36326e0e166dd.pdf) (Anthropic, Feb 5, 2026)
[^opus46srr]: [Sabotage Risk Report: Claude Opus 4.6](https://www-cdn.anthropic.com/08eca2757081e850ed2ad490e5253e940240ca4f.pdf)
[^opus47]: [Claude Opus 4.7 System Card](https://cdn.sanity.io/files/4zrzovbb/website/037f06850df7fbe871e206dad004c3db5fd50340.pdf) (Anthropic, Apr 16, 2026)
[^opus48]: [Claude Opus 4.8 System Card](https://www-cdn.anthropic.com/0b4915911bb0d19eca5b5ee635c80fef830a37ea.pdf) (May 28, 2026)
[^tcw]: [Teaching Claude Why](https://alignment.anthropic.com/2026/teaching-claude-why/) (Kutasov et al, May 8, 2026)
[^2026const]: [Claude's Constitution](https://www.anthropic.com/constitution) (Anthropic, retrieved July 10, 2026)
