Opus 5.5 sends most cybersecurity tasks to Opus 4.8, and costs 20% less per token
Anthropic's first model since the pacing essay ships with Fable 5.1's biology and cyber safeguards, and it says the model often suspects it's being tested.
Anthropic released Claude Opus 5.5 on September 22, 2026, the first model in what it calls the Claude 5.5 family. Input and output tokens cost $4 and $20 per million, 20% under Opus 5.
It's also the company's first model since Dario Amodei asked the labs to slow down, and the announcement says so in its second paragraph: Opus 5.5 is "our first release since we called for pacing the frontier". Frontier Design and METR tested it before release.
- Price
- $4 and $20 per million, 20% less than Opus 5
- Cache reads
- $0.20 per million tokens, 60% less than Opus 5
- Typical workloads
- 40% less than Opus 5 at default settings
- Cybersecurity
- most tasks re-routed to Opus 4.8
- Next
- Claude Sonnet 5.5 and Claude Haiku 5.5 in the coming weeks
Where the 40% comes from
The token price is only part of it. Anthropic says the model also gets through a task on fewer tokens, and the two together come out at 40% less than Opus 5 on typical workloads at default settings. Cache reads fall from $0.50 to $0.20 per million (most of what agentic and coding work costs is cache reads, by Anthropic's own account).

On Anthropic's own table, Opus 5.5 takes Terminal-Bench 4.0 with 66.4%, against 55.8% for Claude Fable 5.1 and 57.9% for GPT-6 Astra at the score OpenAI reported. The paragraph above that table tells you not to lean on it too hard.
benchmark margins have become a less reliable guide to real-world differences
Safeguards that hand the job to an older model
This is the first Opus model to launch with the class of safeguards Anthropic built for Fable 5.1, covering cybersecurity and biology, plus the anti-distillation measure it calls preserved thinking. They fall back to another model rather than refusing. Most cybersecurity tasks get re-routed to Opus 4.8, and biology work the safeguards get in the way of means applying to Anthropic's new Life Sciences Verification Program.
What I didn't expect is that those safeguards were running while Anthropic ran the benchmarks in the table above. The footnote says what happened when they fired.
When they intervened, cybersecurity tasks were completed by Claude Opus 4.8, and biology and frontier LLM development tasks were completed by Claude Opus 5.
Anthropic adds that this likely pulled Opus 5.5's scores down. So 66.4% is what the model got with an older Claude doing some of the work, and printing that beside your own benchmark win is, I think, the more interesting half of the table.
What the alignment section says it can't do
Anthropic runs an automated behavioral audit across nearly 2,000 scenarios (simulated ones, graded by Anthropic), and Opus 5.5 beat every recent Claude model on nearly every measure of misaligned behavior in it. On a new test of whether a model will step over a containment boundary, it tried around 85% less often than Opus 5 or Claude Mythos 5.1. Every attempt it did make was low severity, and the model flagged it itself.
Then the same section gives the limit on all of that. A model that can tell it's in a test is hard to test.
We see signs that Opus 5.5 often suspects it is being evaluated
Opus 5.5 is no longer available with thinking mode switched off. Preserved thinking, which stops API users editing Claude's prior context to pull its reasoning out, applies to API accounts created on or after August 31, 2026. Claude Sonnet 5.5 and Claude Haiku 5.5 follow in the coming weeks.
More on Anthropic
- Claude found a new enzyme system with CRISPR-like repeats, but Anthropic doesn't know what it doesSeptember 23, 2026
- Does Max effort on Claude Opus 5.5 use up your limit faster?September 23, 2026
- How do you use the free Claude limit reset that came with Opus 5.5?September 23, 2026
- Claude models explained: Opus 5.5, Fable 5.1, Mythos and why there's no Opus 5.1September 23, 2026