<model>'s safeguards flagged this message
The message
<model>'s safeguards flagged this messageWhat it means
Anthropic's real-time cyber safeguards on Opus and Sonnet models blocked a request that looked like prohibited or high-risk cybersecurity work. Dual-use work can be unblocked through the Cyber Verification Program; prohibited use can't.
What to do
If the work is legitimate defensive security, have an admin apply to the Cyber Verification Program. If it had nothing to do with security, report it with Send feedback, then edit your message and retry.
Claude Code prints this when Anthropic's real-time cyber safeguards stop a request. Its error reference lists three versions: the model's safeguards flagged this message, flagged this session, or flagged this message for a cybersecurity topic. The model's name goes where the docs write <model>, so on Sonnet 5 you'll see Sonnet 5's name in front.
- Message
- safeguards flagged this message
- Where
- Claude Code
- Cause
- real-time cyber safeguards on Claude Opus and Sonnet models
- Fix for defensive work
- apply to the Cyber Verification Program
- Review time
- within two business days
Why you got it
Anthropic runs classifiers on Opus and Sonnet requests that look for prohibited or high-risk cybersecurity use. They block two kinds of work, and the difference matters for what you can do next.
High Risk Dual use: Cybersecurity activities that have legitimate defensive applications, such as vulnerability exploitation or offensive security tooling development. These are blocked by default, but defensive users can apply for adjustment for legitimate use cases through the Cyber Verification Program described below.
The other kind is prohibited use, like ransomware code or mass data exfiltration. That stays blocked whatever you do. It's worth knowing the check reads more than your last message, too. On Sonnet 5.5, Anthropic says the classifiers review memory, connector content, web search results and files, so a block can come from text you didn't type.
How to fix it
For legitimate defensive work, the route is the Cyber Verification Program. It's free, you apply through the Verification Portal (only authorized admins see the option), and Anthropic aims to email a decision within two business days. You'll need to verify your identity. CVP approval is tied to one organization ID, so a personal workspace won't inherit your team's approval.
If the request had nothing to do with security, report it with Send feedback (Anthropic says those reports are how it narrows the classifiers). Then edit your message and retry.
Messages that look like it
The model's name followed by "can't help with this. Start a new session to continue" is a different error. Claude Code lists it next to the safeguards messages, alongside a line saying the request appears to violate Anthropic's Usage Policy.
On Sonnet 5.5 you often won't see an error at all. A flagged cyber request falls back to Sonnet 5 and re-runs there, with a notice that the model switched (you can turn that off under Switch models when a message is flagged). Biology and distillation blocks don't fall back, though. Those requests are blocked outright.
Anthropic's cyber safeguards article doesn't cover Opus 5.5 or Sonnet 5.5 yet, and it says the Cyber Verification Program will soon expand to include them.