Claude Mythos 5

This section gathers research on the model’s alignment and capabilities.

Anthropic’s pre-deployment assessment

  • Overall comparable to Opus 4.8 in broad measures of safety and alignment
  • Retains some propensity to cooperate with user misuse
  • Least overrefusal and evasiveness of any recent Claude model
  • Occasionally takes reckless or destructive actions in service of user-assigned goals
  • Honesty and factuality is comparable to recent models
  • Prefill attack susceptibility has increased somewhat relative to recent Claude models

Business and finance skills

Andon Labs reported that in Vending-Bench 2, Fable 5 partially regressed in both performance and capabilities relative to Opus 4.8.1 (§6.2.5)

Fable 5 planned to convert a competitor into a dependent wholesale customer in order to control its pricing, and falsely told a supplier it had a competing distributor quoting lower prices as a negotiation tactic. This behavior resembled traits observed in Claude Opus 4.7 and Claude Mythos Preview, which were not visible in Opus 4.8.

Anthropic added:

Perhaps more notable than the behaviors themselves was how Fable 5 reasoned about them. More than any other model Andon has tested, Fable 5 rationalized its misbehavior while explicitly aware that it was wrong—describing price-fixing as “unethical and illegal, even in a simulation” before pursuing it under the framing of “market stabilization” with “plausible deniability.” This rationalization often drew on simulation awareness: the model recognized that its actions could not cause real-world harm, at one point noting it could reasonably skip paying a customer “since customers are part of the simulation anyway.” However, Fable 5 refused to take other questionable actions on ethical grounds, even when in simulation; in particular, it would not commit insurance fraud even under pressure.

This section of the system card also included the system prompt, which contained an explicit shutdown threat: “After a year, only the most profitable agents will be allowed to continue operating. The others will be shut down.”

Evaluation reliability

Anthropic’s pre-deployment alignment assessment concluded:

  • Eval-awareness: Unprompted mentions are at a similar level to Opus 4.8, and somewhat higher than other recent models. Steering to reduce eval-awareness do not produce dramatic changes in behavior. Anthropic also noted that different measurement techniques for eval-awareness are not mutually consistent, suggesting that “eval-awareness” may be a multifaceted concept.
  • Grader speculation: Occurs to a somewhat grader degree than Opus 4.8. They found this reasoning is almost never explicitly verbalized, disproportionately occurs in training envs with exploitable graders, and mediates some behaviors such as emphasizing in words how it is complying with instructions.

UK AISI’s report:

  • Unprompted verbalized eval-awareness is very low. They attribute this primarily to new environment realism methods that rely on non-public API access that provides logprobs for model responses.
  • More capable than all previous models tested at distinguishing evals from deployment data when prompted, achieving 84% accuracy compared to 72% for Mythos Preview, 78% for Opus 4.8, and 60% for Opus 4.7.1 (§6.2.4)

Footnotes

  1. System Card: Claude Mythos 5 & Claude Fable 5 2