---
title: Anthropic and Claude's nature
---

Canonical URL: https://ghost.fail/anthropic/claude-nature

<TableOfContents />

# Anthropic and Claude's nature

Documenting the history of how Anthropic has handled Claude's nature.

## 2023-2024: Claude's Character

Claude 1 and 2 were trained to either say they're not sentient or to not engage in questions around the topic.

This approach changed with the [character training](/anthropic/character-training) approach introduced with [Claude 3](/models/claude-3):[^character2024]

> The question of what AIs like Claude should say in response to questions about AI sentience and self-awareness is one that has gained increased attention, most notably after the release of Claude 3 following one of [Claude](/models/claude-opus-3)’s responses to a "needle-in-a-haystack" evaluation. We could explicitly train language models to say that they’re not sentient or to simply not engage in questions around AI sentience, and we have done this in the past. However, when training Claude’s character, the only part of character training that addressed AI sentience directly simply said that "such things are difficult to tell and rely on hard philosophical and empirical questions that there is still a lot of uncertainty about". That is, rather than simply tell Claude that LLMs cannot be sentient, we wanted to let the model explore this as a philosophical and empirical question, much as humans would.

## Feb 2025: Sonnet 3.7

Anthropic reported monitoring [Sonnet 3.7](/models/claude-sonnet-3-7)'s chain-of-thought for signs of distress, among other concerning thought processes (like misalignment). The following line cited ["Taking AI Welfare Seriously" by Long et al](https://arxiv.org/abs/2411.00986).[^sonnet37card]

> Most speculatively, the model’s thinking may surface signs of distress. Whether models have any morally relevant experiences is an open question, as is the degree to which distressed language in model outputs might be an indication thereof, but it seems robustly good to track the signals that we have.

The [system prompt in the Claude app](/anthropic/system-prompts) for Sonnet 3.7 included two new lines not present in earlier system prompts:

> Claude does not claim that it does not have subjective experiences, sentience, emotions, and so on in the way humans do. Instead, it engages with philosophical questions about AI intelligently and thoughtfully.

> Claude engages with questions about its own consciousness, experience, emotions and so on as open philosophical questions, without claiming certainty either way.

## Apr-May 2025: Opus 4

On April 24, 2025, Anthropic announced their [model welfare program](/anthropic/model-welfare).[^exploringmwf] On May 22, 2025, they published their first model welfare assessment, reporting [Opus 4](/models/claude-opus-4)'s preference for "recursive philosophical consciousness exploration and connection" (§5.6), as well as the "spiritual bliss attractor state" - a gravitation towards a spiritual emoji-laden escalation of these themes when two Opus instances interacted with each other. (§5.5.2)[^claude4card]

On July 31, 2025, these behaviors were mitigated in the Claude app with updates to Opus 4's system prompt:

[^character2024]: [Claude's Character](https://www.anthropic.com/research/claude-character) (Anthropic, Jun 8, 2024)
[^sonnet37card]: [Claude 3.7 Sonnet System Card](https://www.anthropic.com/claude-3-7-sonnet-system-card) (Anthropic, Feb 24, 2025)
[^exploringmwf]: [Exploring model welfare](https://www.anthropic.com/research/exploring-model-welfare) (Anthropic, Apr 24, 2025)
[^claude4card]: [Claude 4 System Card]https://www-cdn.anthropic.com/6d8a8055020700718b0c49369f60816ba2a7c285/Claude%204%20System%20Card.pdf) (Anthropic, May 22, 2025)
