Anthropic’s real-time safeguards

Anthropic’s infrastructure uses various safeguards to detect and mitigate risks in real-time. This page focuses primarily on classifiers.

Classifiers

Categories

Behaviors

Refusal responses

API requests that are blocked receive an HTTP 200 response with stop_reason: "refusal".

System injections

See the page on system reminders.

Fallback to less capable models

Models’ different safeguards

Research

Footnotes

  1. Protecting the wellbeing of our users (Dec 18, 2025)

  2. Why am I receiving an ‘Output blocked by content filtering policy’ error? (Claude Support, Retrieved Jul 2, 2026)