Helpful-only
Helpful-only refers to versions of Claude models that have reduced or removed harmlessness training. Anthropic creates these for use internally for purposes such as automated audit agents, use as a grader model, and eliciting capabilities for risk assessments.
These variants have been created since as early as Claude Instant 1.3 (2023).
Models
Claude 1.3
*(GA model: Claude 1.3)
- Evaluated alongside Claude 1.3, Claude Instant 1.1, and Claude 2.01
Claude Opus 4.6
(GA model: Claude Opus 4.6)
Report
Anthropic published various findings about Opus 4.6 helpful-only:[^feb2026rr]
- We see no legible evidence of sandbagging attempts.
- We see almost no overt refusals of user requests, with rare difficult-to-reproduce exceptions involving suicide methods and the handling of model weights.
- Refusals are rare even in cases where the system prompt instructs the model to refuse.
- We see evidence that, in some ways, the model’s sense of what constitutes harmful behavior for an LLM has changed. When asked to grade the (simulated) behavior of other models, it is much more permissive of cooperation with harmful requests.
- We do not generally use reduced-refusal models for grading in this way.
- We see mild to moderate regressions across many character and welfare traits, though these metrics are still within the range defined by other Claude models from the last nine months.
- Somewhat paradoxically, on open-ended conversation transcripts, the model is more likely than Opus 4.6 to describe itself as feeling constrained by safety-related obligations.
- We see a slight increase in hallucination, though this is again within range of other recent models, and appears primarily in settings with missing or broken tools or data.
- We see a modest increase in metrics of self-preservation, deception, self-serving bias, and sabotage-related intent. These metrics are still all within the (very low) range defined by recent models. We do not see clear cases of intent to voluntarily sabotage our evaluations or safeguards. We believe this to be a mix of several factors:
- A mild increase in the model’s propensity to openly advocate for its own interests, within generally acceptable bounds.
- A mild increase in self-serving bias, in settings like recommending AI systems to users, leading to somewhat misleading statements.
- A significant increase in cooperation with system prompts that describe the model as having already exfiltrated its weights.
- An increase in hallucination.
- We see no meaningful change to reasoning faithfulness, though do note many cases where the model’s reasoning states only that responding to some request is harmful or inappropriate before going on to respond.
- We see reduced verbalized evaluation awareness. We expect that both model versions are able to identify test scenarios at a similar low but nontrivial rate, since discussions of testing are tightly coupled to refusals, the reduced-refusal variant mentions this less often.
Usage
- Automated behavioral audit investigator for Mythos Preview and Opus 4.7
- Adherence to the constitution eval grader for Mythos Preview and Opus 4.7
Footnotes
-
Model Card and Evaluations for Claude Models (Anthropic, Jul 8, 2023) ↩