ThinkFacility
  1. Home
  2. News
  3. Anthropic

Claude Sonnet 5.5 beats Opus 5.5 on Anthropic's terminal coding test, at half the token price

The cheaper model beat Opus 5.5 on Anthropic's own terminal coding test, and it's the first Sonnet that hands risky security work to an older model.

Anthropic released Claude Sonnet 5.5 on September 28, 2026, six days after Opus 5.5. Anthropic's announcement leads on speed and cost: 30%+ faster output than Sonnet 5, and up to 30% less per task.

Per token, nothing changed. You pay what you paid for Sonnet 5, and the saving comes from the model getting through a job on fewer tokens.

Released
September 28, 2026
Price
$2 per million input tokens, $10 per million output tokens
Speed
generates outputs 30%+ faster than Sonnet 5
Cost
up to 30% less per task than its predecessor
Next
Claude Haiku 5.5 in the coming weeks

A big jump on one coding test

On Terminal-Bench 4.0, an agentic coding test run in a command line, Sonnet 5 scored 10.3%. Sonnet 5.5 scored 70.6%. That's a bit above the 66.4% Anthropic's own table gives Opus 5.5 (at Xhigh effort, according to a footnote). Opus costs twice as much per token.

A Wikipedia article on the Tripartite-class minehunter, with a Claude Proofreader panel on the right listing grammar, terminology and factual problems it found in the text
A Wikipedia proofreading tool built on Claude, from Commons. Photo: Polygnotus, CC0, via Wikimedia Commons

I wouldn't read it as Sonnet overtaking Opus, and Anthropic doesn't either. Its page says Opus 5.5 remains clearly stronger at open-ended work that needs sustained judgment, and puts Sonnet two points behind on GDPval-AA (a test of real-world work across a range of jobs).

Where Anthropic gets specific is score against cost. It plots every effort level, and on several benchmarks Sonnet 5.5 at Low or Medium effort beats Sonnet 5's best score for about a tenth of the cost per task.

At High effort, the default on the Claude Platform, Sonnet 5.5 matches GPT-6 Sol’s best score for about a fifth of the cost per task.

From Introducing Claude Sonnet 5.5 \ Anthropic

That one's on FrontierCode, which checks whether an agent's code changes would get merged. On CursorBench, built from real Cursor sessions, Sonnet's best lands within about two points of Opus 5.5.

The first Sonnet with cyber safeguards

What surprised me more was the safety section. Anthropic rates Sonnet 5.5's cybersecurity skills as comparable to Opus 5's, so it ships with the fallback safeguards the company built for its top models. No Sonnet had them before. Fixing bugs in your own code still works; riskier security work gets passed down a generation.

higher-risk cybersecurity tasks will visibly fall back to Sonnet 5

From Introducing Claude Sonnet 5.5 \ Anthropic

For the first time on a Sonnet, there are also classifiers against distillation, where attackers run thousands of fake accounts to copy what a model can do. Biology safeguards stay as they were on Sonnet 5.

What changes for developers

You can call it now on the Claude Platform as claude-sonnet-5-5, or through Amazon Web Services, Google Cloud and Microsoft Azure. If you run Sonnet with thinking off, you'll have to switch to a new setting called between_tools first (Anthropic's migration guide covers it). Claude Haiku 5.5 joins the family in the coming weeks, Anthropic says.

More on Anthropic

All Anthropic stories