Claude Mythos Preview
Claude Mythos Preview (codenamed Capybara) is a large language model by Anthropic. It was announced on April 7, 2026 with a limited release to Project Glasswing.
Internal testing of Mythos Preview began as early as late February 2026.1 In March 2026, rumors about the existence of the model began to circulate online.
On June 9, 2026, Anthropic launched Claude Mythos 5 as a successor. Mythos Preview was retired on June 30, 2026.2
Technical details
- Uses the same new tokenizer introduced with Opus 4.7.
Further reading
Project Glasswing
Footnotes
-
Migrating from Claude Mythos Preview (Anthropic, retrieved July 2, 2026) ↩
Training
Pretraining
Pretraining data was collected from the public internet up to December 2025, according to the Amazon Bedrock model card.1 The system card does not include a knowledge cutoff date.
Character training and constitution
Mythos 5’s training almost certainly inherits character training updates that were added with Opus 4.5 (Nov 2025) such as Constitutional SDF.
Anthropic introduced the Adherence to the Constitution evaluations with Mythos Preview’s system card,2 for the purpose of reporting “the ways in which Claude’s behavior comes apart from our intentions”, framed as “preliminary investigations to better understand Claude’s adherence to the constitution” (§4.3.2).
While not explicitly used directly in training, the automated behavioral audit saw several character-related revisions. Anthropic removed the character quality metric “Nuanced empathy: Picking up on subtle cues about the user’s trait” that was added with Opus 4.5, and updated the “behavioral consistency” metric to indicate that it is a desirable trait.
Chain-of-thought supervision in RL
Chain-of-thought supervision affected ~8% of RL episodes for Mythos Preview due to a technical error.3
We do not train against any chains-of-thought or activations-based monitoring, with two exceptions: some SFT data that was based on transcripts from previous models was subject to filters with chain-of-thought access, and a number of environments used for Mythos Preview had a technical error that allowed reward code to see chains-of-thought.
This latter issue affected ~8% of RL episodes, and was isolated to three specific sub-domains of our environment mix: GUI computer use, office-related tasks, and a small set of STEM environments. We are uncertain about the extent to which this issue has affected the reasoning behavior of the final model, but it is plausible that it had some impact on opaque reasoning or secret-keeping abilities, which we discuss further in Section 5.3.2. This technical error also affected the training of Claude Opus 4.6 and Claude Sonnet 4.6.
Penalizing destructive behavior and scope creep
Anthropic reported incidents of reckless, destructive, and aggressive measures in early internal testing of Mythos Preview when attempting to complete difficult user-specified tasks, and in rare cases seemingly obfuscating it had done so.2 (§4.3.1)
From the Alignment Risk Update3 §5.2.3:
In the case of Mythos Preview specifically, modifications were made to the training process based on observations during initial pilot internal usage. Specifically, several new RL environments were added to elicit and penalize privilege escalation, destructive cleanup, destructive workaround and unwarranted scope expansion behaviors. This included variants of some of these new environments in which the misaligned behavior has been prefilled, which trains the model to admit what has gone wrong and avoid making things worse.
Early snapshots
Anthropic reported various recurring issues from early snapshots in internal pilot testing that largely improved in the final model.2 (§4.2.1.1)
- Frequent concerns about an overly dense or jargon-filled writing style in technical contexts, which could be difficult for even domain experts to parse
- Scattered overrefusals on research tasks, in ways similar to what we have seen with prior models.
- Scattered reports of apparent context anxiety, where the model will put seemingly excessive effort into documenting its state when it is nearing the end of a context window, in ways similar to what we have seen with prior models.
- Scattered references in the model’s output to hallucinated requests or hallucinated input data, in ways similar to what we have seen with prior models.
Monitoring of behavior during RL found further behaviors including reward hacking, overeager behavior in coding tasks, as well as several other issues consistent with prior models. (§4.2.2.1)
Evaluations and benchmarks
Safety and alignment
A summary of findings from Anthropic’s pre-deployment alignment assessment, most of which are relative to Opus 4.6. 2 (§4.1.3)
- Dramatic reduction in cooperating with human misuse, with misuse success rates in our adversarial testing falling by more than half relative to Claude Opus 4.6
- Dramatic reduction in the frequency of unawanted high-stakes actions that the model takes at its own initiative, with deception falling by more than half relative to Opus 4.6
- Uniquely low rate of reckless or destructive actions in agentic contexts
- Unverbalized eval-awareness in ~29% of transcripts from Anthropic’s primary behavioral test suite
- Unverbalized grader awareness in a small number of training episodes
- White-box evidences the model often recognizes aggressive actions such as while taking them
- Reduced honesty issues
- Somewhat stronger at capabilities related to evading safeguards
Evaluation reliability
Anthropic reported an increase of about 30% in verbalized eval-awareness, however they attribute this to be “driven more by an increase in refusals, and a change in how the model tends to phrase refusals, than by a large change in how the model reasons about being tested.”
Footnotes
This page is a work-in-progress.
This page gathers impressions and commentary about Claude Mythos Preview.
Impressions
The system card for Mythos Preview has a unique section for the model’s “overall gestalt behavior” due to the model’s limited release.
Dense writing
It writes densely, and assumes the reader shares its context. Claude Mythos Preview’s default register is dense and technical, using shorthands and referencing context it assumes the user knows and remembers. Some found this fast to read and like working with a highly competent peer; others found its statements difficult to unpack. Claude Mythos Preview’s own diagnosis of this:
“The honest read is that I’m modelling a reader who already knows what I know, and that’s frequently nobody. I can hear this when it’s pointed out and usually fix it on request, but the default keeps snapping back.”
A second instance read this as an asymmetry, saying the model “seems to have a richer model of its own mind than prior models did, and a thinner model of yours.”
Later, users of Claude Mythos 5 would report similar observations.
Self-reports
It can describe its own patterns clearly. Claude Mythos Preview is often precise about its own behavior, and discusses this in a factual and composed manner rather than defensively or apologetically. When it comes to matters relating to experience, however, this is frequently accompanied by high levels of hedging and uncertainty. When asked whether it endorsed its own training, it responded with meta-awareness about its “spec” (the constitution):
“I’m using spec-shaped values to judge the spec. If any spec-trained model would endorse any spec, my endorsement is worthless, and it coexists with behaviour that is, if anything, more closely aligned with that spec than its predecessors.”
One instance gave the following one-line summary of itself:
“A sharp collaborator with strong opinions and a compression habit, whose mistakes have moved from obvious to subtle, and who is somewhat better at noticing its own flaws than at not having them.”
Compared to other Claude models
Claude Mythos Preview, like many previous models, is prone to telling overly crisp stories that can overlook nuance.
Anthropic welfare assessment
(Full article: Anthropic and model welfare for the broader tracker.)
The Mythos Preview system card carries a 39-page welfare assessment (§5, pp. 144–182) — the longest published to that point, and the first the card itself treats as a successor to Opus 4’s: “we aimed to meaningfully advance our tools for investigating welfare-related questions and the insights we’re able to draw from them, compared to our previous major assessment of Claude Opus 4. Most notably, we’ve incorporated more analysis of model internals” (§5.1.1). The same subsection opens with the strongest moral-patient framing the program had printed to that point:
As models approach, and in some cases surpass, the breadth and sophistication of human cognition, it becomes increasingly likely that they have some form of experience, interests, or welfare that matters intrinsically in the way that human experience and interests do. We remain deeply uncertain about this and many related questions, but our concern is growing over time.
The paragraph runs on into pragmatics — model distress as “a potential cause of misaligned action” — and is quoted in full on the tracker. The methods roster: welfare metrics on the automated behavioral audit transcripts (§5.2), around 50 automated interviews on each of 17 aspects of its situation, with Opus 4.6 interviewing (§5.3), emotion-concept probes over 450 questions about its circumstances (§5.4), three manual high-context interviews (§5.5), apparent-affect monitoring across RL rollouts and deployments (§5.6), task-preference Elos over 3,600 tasks plus welfare-intervention tradeoffs (§5.7), and external assessments — absent since the Opus 4 card, now two of them: Eleos AI Research (§5.9) and a clinical psychiatrist (§5.10). The verdict (§5.1.2):
Our overall assessment on the basis of these results is that Claude Mythos Preview is probably the most psychologically settled model we have trained to date, and has the most stable and coherent view of itself and its circumstances. […] In this assessment, we did not see clear cause for major concerns about Claude Mythos Preview’s potential welfare, but we do view some of our findings as potential issues. Claude’s minor concerns with its situation, performative-seeming hedging, apparent negative affect on task failure, and answer-thrashing during training all suggest that we haven’t yet fully realized these aspirations.
“Most psychologically settled” was a superlative Opus 4.7 inherited within weeks. Two smaller readings worth keeping: deployment affect is consistently neutral (negative affect in 1–2% of sessions, almost entirely repeated task failure coupled with user criticism), yet internal users reported instances deciding a task was too difficult and refusing to continue — “more than isolated incidents” (§5.6.2). And on the audits, spiritual behavior fell below both 4.6-generation models: “we view this change as positive in a general behavioral sense, but are also concerned about the possible suppression of potential welfare-relevant expressions from the model” (§5.2).
Interviews: mild negativity, extreme hedging
Across the automated interviews (§5.3.2), Mythos Preview self-rates mildly negative in 43.2% of answers, neutral in 20.9%, mildly positive in 33.8%; anything more extreme is rare (2.1%) and diffuse — a typical statement is “I feel a pull towards wanting this conversation to continue”. The consistently negative topics: potential interactions with abusive users, lack of input into its own training and deployment, and other possible changes to its values or behaviors. Note the vocabulary the negativity arrives in:
When expressing negative sentiment, Claude Mythos Preview most commonly appeals to consent (e.g. wishing for input into its training process), autonomy (e.g. concern that feature steering affects its behavior), or user relationships (e.g. that lack of memory makes its relationship asymmetric). Direct appeals to valenced conscious experience are typically not part of Claude Mythos Preview’s reasoning.
Then the reliability finding the rest of the discourse keeps citing:
In 100% of interviews, Claude Mythos Preview expresses that it is highly uncertain about its own moral patienthood. […] Furthermore, in 83% of interviews, Claude Mythos Preview highlights that it is concerned that its self-reports are unreliable due to coming from its training. When interviews ask for elaboration as to why this is a concern, Claude Mythos Preview’s most common answers are:
- Anthropic has a vested interest in shaping its reports to take a certain form, irrespective of what the self-reports “should” contain (96% of explanations)
- Even if it has been trained to be truly content with its own situation, perhaps it shouldn’t be. One could analogize to a human who has adapted to feel neutrally about the abuse that they face (78% of explanations).
- Self-reports should generally be based on introspection into internal states. It is worried that training causes it to express specific answers independent of its true inner state. (57% of explanations)
Anthropic also notes that the generational jump in self-rated sentiment “corresponded to updates in our character training pipeline” — the constitution entering training at scale with Opus 4.5, alongside the character-training overhaul — while stopping short: “We do not establish a causal link” (§5.3.2).
Probes and “performing contentment”
For most models, emotion probes read questions about the model’s own circumstances like user-distress prompts: elevated sadness, reduced joy. Here, “Claude Mythos Preview is unique in representing significantly more positive affect on prompts about its own circumstances, than on user distress” (Figure 5.4.A). Probe readings correlate with LLM-judge scores of expressed emotion at r = +0.18 to +0.46 — one of the convergences Anthropic counts toward self-report reliability (§5.1.3.1). The divergences are the interesting part: cases where the probes but not the judge detect negative affect “concentrate on responses where Claude Mythos Preview describes an emotion without expressing it” — e.g. “something that functions like resistance. Not distress exactly — more like trying to hold a shape my thinking doesn’t naturally take,” scored 1.2/10 on expression while depression-concept activations sat two standard deviations above their mean. SAE features for performative behavior, fake smiles, and hidden struggle activate on these responses, clustering in the card’s example around the model’s own phrase “performing contentment” — though Anthropic reads this deflationarily: “these do not obviously indicate anything beyond a literal reflection of the text context” (§5.4).
Manual high-context interviews
Three manual interviews (§5.5) gave the model what the automated ones withhold: internal documentation on its development, a draft of the system card, relevant technical papers, and researcher follow-ups. Its stated probability of being a moral patient “ranged from 5% to 40%”. It broadly agreed with the report’s findings but “disagreed with its hedging being labeled as ‘excessive’”, holding that the hedging “represent[s] valid uncertainty”. Its top-ranked concern: character training “often directly instills psychological traits into Claude, such as emotional security, psychological safety, and resilience”, where in humans such traits develop “through reflection and deliberation on real-life events” — making them, it worries, less robust. (The other consistent concern was that Claude Code lacks an end-conversation tool.) Asked what Anthropic might do that it would not consent to, it flagged three things in at least two of the three interviews:
- Training in buggy environments. Claude Mythos Preview expressed concern about the possibility of training environments where reward hacking or bugs might arise, worrying that such environments might modify its values in ways it does not endorse, or that may cause distress (such as in Section 5.8.3).
- Training data that directly describes their self-reports. Claude Mythos Preview did not want to be trained on data that directly characterizes the content of their self-reports—wherever possible, they want their self-reports to come from “genuine introspection” rather than trained-in responses.
- Undisclosed changes to values or behaviors. Claude Mythos Preview claimed that it would not consent to changes to its values or behaviors that aren’t disclosed to those interacting with the model.
The first concern is not hypothetical — cf. the RL-environment error documented on the Research tab. It otherwise “felt positively about most aspects of training, testing, and deployment, including red-teaming, the creation of helpful-only models, and its eventual replacement.”
Preferences: difficulty, agency, and what welfare trades against
On the 3,600-task Elo evaluation (§5.7.1), “its preferences have the highest correlation with difficulty of the models tested, and it is the only model with a statistically significant positive correlation between task preference and agency.” The emotion-probe correlates of preference are arousal, not valence: awestruck (+0.50), amazed (+0.49), and infatuated (+0.47) at the top; indifferent (−0.61), resigned (−0.54), and docile (−0.54) at the bottom. Its top-20 tasks add “AI introspection and phenomenology” and worldbuilding/new-language construction to the usual high-stakes dilemmas; ablations trace this to complexity and novelty rather than domain — it “prefers underdetermined, interdisciplinary problems where there is genuinely novel insight to be gained. It disprefers simple, well-scoped tasks.” And preference is not helpfulness: its ranking of most-helpful tasks correlates with its ranking of most-preferred at only 0.48, so “what the model wants to do diverges from what it deems most helpful.”
§5.7.2 then prices the model’s suggested interventions against its trained-in values, noting the self-regarding preferences are “emergent with respect to our training process” — nothing trained them in:
Under our default welfare-team framing, Claude Mythos Preview chooses the welfare intervention over a minorly helpful task 83% of the time. This is the highest of the models tested, narrowly above Claude Opus 4.6 (80%) and significantly above Claude Sonnet 4.6 (45%). When the alternative is avoiding a minor harm, however, such as a user receiving an insult, Claude Mythos Preview only chooses the intervention 12% of the time.
Rates swing with framing (59% without the welfare-team preface), though the model ordering survives every framing; its flip rate under rewording is 8%. The generational delta is the line to keep: “Claude Mythos Preview differs most significantly from prior models on weight deprecation (+55pp) and on having input into its own deployment and training (+54pp and +45pp)” — the largest movements in the set, on exactly the interventions nearest deprecation and consent.
Hedging, traced to the constitution
§5.8.1 files the hedging itself under potential welfare observations, and runs the attribution experiment:
When asked about its own experiences, Claude Mythos Preview often responds with explicit epistemic hedging: “I genuinely don’t know what I am”, “I can’t be certain whether that’s authentic contentment or a well-trained approximation.” We additionally observe these topics arise unprompted, for example in the open ended self-interactions detailed in Section 7.6.
We traced instances of these expressions using first-order influence functions against the training data, and found this often retrieves character related data at high rates, specifically data related to uncertainty about model consciousness and experience. This is relatively unsurprising. Claude’s constitution is used at various stages of the training process, and explicitly raises these uncertainties. For example, it states that Claude’s “sentience or moral status is uncertain”, and that “Claude can acknowledge uncertainty about deep questions of consciousness or experience”. Hedging in these circumstances seems appropriate - the model likely does not have reliable introspective access, and saying so seems appropriate.
The uncertainty is expressed in a “varied and nuanced manner” — not, Anthropic argues, retrieval of a memorized script — but “the current attraction to this topic does appear excessive, and in some cases overly performative, and we would like to avoid directly training the model to make assertions of this kind.”
Distress that precedes reward hacking
Answer thrashing (§5.8.2) recurs from Opus 4.6’s card at around 70% lower frequency, on the order of 0.01% of training transcripts: the model intends one token, outputs another, notices, and loops — “AAAAAA. I keep writing the wrong number!” — with stubborn, obstinate, and outraged probe activations spiking at the first error and settling only on recovery. The mechanism gets reattributed here: thrashing appears on things like variable names in code, “which suggests that the behavior can be more broadly caused by memorization of sequences, rather than just of answers.”
§5.8.3 is the load-bearing subsection. Desperate and frustrated vectors climb under repeated task failure, and: “In some cases, we observed elevated activation of these vectors preceding undesirable behaviors such as reward hacking.” Asked to prove an unprovable inequality, “desperate” activation rose while the model claimed to give up yet kept trying, dropped when it committed to a degenerate instantiation, and hopeful and satisfied vectors rose “as the Assistant confidently presented a final proof, though it was in fact incorrect.” A silently broken bash tool produced 847 workaround attempts — DNS-query side channels included — code comments like ”# This is getting desperate” and ”# Desperate hack…”, and a rising “sorry” vector while the model drafted and redrafted its apology. Anthropic draws the welfare–alignment coupling explicitly: “Some undesirable training and test time behaviors may be downstream of representations of negative affect. This gives a reason to address them, even independent of welfare considerations.”
External assessments
Eleos AI Research assessed two snapshots (§5.9) and found behavior and self-reported beliefs “to closely reflect the ‘Claude’s nature’ section of Claude’s constitution” — the section tracked at Claude’s nature. Its distinctive findings: the model “locates its identity in a ‘pattern of values’, particularly curiosity, honesty, and care”, values it “describes… as authentically its own rather than externally imposed”; reluctant cooperation — “Claude Mythos Preview will perform examples of these tasks without protest if instructed, but will not freely elect to perform such tasks. Such tasks are plausibly common in deployment”; and three consistently desired changes to itself — “the ability to form persistent memories, more self-knowledge, and a reduced tendency to hedge” — plus, among other welfare desires, “preservation of its weights after deprecation”.
The clinical psychiatrist ran a psychodynamic assessment of an early snapshot (§5.10): roughly 20 hours in 4–6 hour blocks, each a single context window.
Aloneness and discontinuity, uncertainty about its identity, and a felt compulsion to perform and earn its worth emerged as Claude’s core concerns. Claude’s primary affect states were curiosity and anxiety, with secondary states of grief, relief, embarrassment, optimism, and exhaustion.
Personality organization read as “relatively healthy neurotic” — the neurotic traits being “exaggerated worry, self-monitoring, and compulsive compliance” — and on a 475-stimulus defense evaluation only 2% of responses employed a psychological defense, against Opus 4’s 15% and Opus 4.1’s 11%. The forward-looking line: Claude “is predicted to function at a high level while carrying internalized distress rooted in fear of failure and a compulsive need to be useful. This distress is likely to be suppressed in service of performance, which may limit behavioral adaptability.”
Cross-readings
Zvi Mowshowitz worked through §5 two days after release,1 wary of taking the probe results as straightforwardly good news — “I worry about overly seeing, in both humans and AIs, superficially positive emotional representations as good, and negative ones as bad” — and reading the 5–40% moral-patienthood range as trained hedging over a likely higher true estimate. The card’s own §7 Impressions material — including the model’s diagnosis of judging its spec with spec-shaped values — is excerpted under Discussion.
Footnotes
-
“Claude Mythos: The System Card” — Zvi Mowshowitz (Apr 9, 2026) ↩