Anthropic Introduces 'Claude Sonnet 5.5,' a Cheaper Model That Catches the Flagship by Finishing Sooner, Not by Thinking Harder

For most of the past few years, a new model announcement meant a higher score and a higher bill. The interesting part of the past month however, the middle of the catalog has started to close on the top, and the labs are explaining that less as a breakthrough in raw ability than as a change in how the model spends its effort.

'Claude Sonnet 5.5' arrived on September 28, six days after Opus 5.5. 

Anthropic did not present it as a replacement for the flagship. It presented it as the model users leave on the everyday jobs: a bug with a clear reproduction, a feature with a spec, a slide deck that has to follow a template, a spreadsheet that has to come back clean. 

The list price is unchanged from Sonnet 5, $2 per million input tokens and $10 per million output, half of Opus. 

What changed, on Anthropic's account, is the path the model takes to an answer. 

It reads a codebase faster, batches its tool calls instead of stepping through them one at a time, and stops sooner. 

In their tests that meant more than 30% faster output and up to 30% less spent per task, not because the meter got cheaper but because the job used fewer tokens. 

Early users described the same thing in plainer language. CodeRabbit said the old habit of reaching for web search too often was gone. Balyasny, on a private set of finance tasks, saw token use fall from about 497,000 per answer to about 121,000.

Slack kept its prompts the same and got better results in fewer steps. 

The other lever is effort, which Anthropic now exposes as a dial rather than a fixed personality. 

At the low and medium settings the model answers faster and checks less. At max it reasons longer and reviews its own work. The claim is that the low settings already beat the best Sonnet 5 could do, at about a tenth of the cost on several of their tests, and that the high settings land near Opus on a lot of scored work. 

That is a different product story from "we trained a smarter model." 

Instead, it is closer to "we trained a model that knows when a task is finished."

Opus is still described, by Anthropic and by outside testers, as the one that holds up when the job is open-ended and the right next step is not obvious. Sonnet is the one you point at a target and a way to check it.

Independent testing from Artificial Analysis found that at maximum effort, Sonnet 5.5 used more output tokens than any model it had measured, about 60% more than Opus 5.5, and that the all-in cost of those long runs was higher for Sonnet than for Opus even at half the token price. 

Matching the flagship, in other words, can mean working harder rather than working smarter. 

The saving shows up when you do not ask it to match the flagship. 

That is the implication buried in the pricing table. 

A company that routes every agent to max effort may not see the discount Anthropic is advertising. A company that reserves the long reasoning for the ambiguous work, and runs the rest cooler, will.

The practical effect is a split in how these systems get deployed. 

Opus, and the more expensive Claude Fable 5.1 above it on the price list, stay in the path where a wrong judgment is costly and the task has no clean definition of done. 

Sonnet becomes the default behind the features people actually hit all day: support replies, code review on ordinary pull requests, first drafts of documents, interface polish. 

Zendesk reported tickets moving about 20% faster on its support cases. Box said the model rechecked figures against source documents and caught errors the previous Sonnet missed. Base44, building apps, saw quality level with an older Opus at roughly half the iterations. 

None of that is a claim about a new ceiling. 

It is a claim about less waste between the question and a usable result.

Against the other labs the same pattern holds, with one exception. 

On price Sonnet 5.5 sits with GPT-6 Sol. On the work Anthropic and the outside evaluators published, it is ahead of that model by a wide margin on knowledge tasks and close to or ahead of Opus on terminal-style coding. Google's Gemini 4 Argon still led the Vals Index in early October, a GDP-weighted mix of finance, coding, legal, and tax work, with Sonnet a fraction behind Opus and both of them ahead of Fable 5.1. The mid-tier label is starting to describe a billing tier more than a capability tier. If that sticks, the procurement question stops being which lab has the smartest model and becomes which jobs are worth the slow, expensive pass. 

There is a constraint on how far that routing can go. Sonnet 5.5 is the first Sonnet with the cyber safeguards used on Anthropic's more capable models, and higher-risk requests in that area fall back to Sonnet 5. Routine development is supposed to be unaffected. Biology safeguards are unchanged. The fallback is a quiet tax. A workflow that looks uniform can, on a subset of tasks, be answered by an older model without the user having asked for that. Teams that care about predictability will have to decide whether that is an acceptable trade for keeping the cheaper model in the loop.

Haiku 5.5 is still promised for the high-volume end. If it arrives in the same shape, Anthropic will have three rungs that differ less in what they can do than in how long they are willing to think and what that thinking costs.

Published