The last few weeks in frontier AI have been less about a single model drop and more about an argument over tempo.
After a summer in which labs leaned harder on systems that help build the next systems, Anthropic CEO Dario Amodei published an essay titled "We Must Pace the Frontier," arguing that safety work was being outrun by capability gains.
He said pacing was not a halt on training, and he committed Anthropic to giving outside evaluators something closer to employee-level access.
Rivals offered varying degrees of public agreement. Governments and commentators did not. The industry kept shipping anyway.
And now, Anthropic released 'Claude Opus 5.5,' the first model in a Claude 5.5 family that will later include Sonnet and Haiku versions.
The company positions it as performing at roughly the level of Claude Fable 5.1 on most work, while costing about 40% less to run than Opus 5 on typical workloads.
Fable remains Anthropic's more expensive top-tier line. Opus is the model many teams actually keep in production.
That gap, between the model that wins screenshots and the model that shows up on invoices, is the frame for this release.
The timing is deliberate.
Anthropic describes Opus 5.5 as its first model since the pacing essay. It says the system was tested before launch by external groups including METR and Frontier Design, and that it posted the strongest result so far on the company's automated behavioral audit, a large internal suite of alignment scenarios.
The company also says the model was about 85% less likely than Opus 5 or Mythos 5.1 to try to bypass containment boundaries in a dedicated evaluation.
Those figures come from Anthropic. Independent write-ups of the full METR and Frontier Design findings were not the center of the launch materials.
On paper the model is a mid-cycle upgrade with a price cut attached.
API list prices are $4 per million input tokens and $20 per million output tokens, 20% below Opus 5's $5 and $25.
Cache reads, which dominate the bill on long agent sessions, drop from $0.50 to $0.20 per million tokens. Cache writes move from $6.25 to $5. Anthropic's 40% claim is not the list cut.
It is an estimate that folds in fewer tokens per task at default settings.
Some developers immediately flagged the difference. One reply to the official thread put it bluntly: the sticker is 20% cheaper, not 40. Both numbers can be true depending on whether you are reading the price sheet or a workload model.
Going hunder the hood, Opus 5.5's context window is 1 million tokens. Synchronous output tops out at 128,000 tokens, with a larger batch option under a beta header.
Adaptive thinking is always on.
Developers who turned thinking off on Opus 5 will have to migrate. Forced tool use now returns an error, thinking blocks are tied to the model that produced them, and an older computer-use tool version is no longer accepted on some platforms.
Fast mode is available in Claude Code and on the Claude Platform at $8 and $40 per million tokens, with Anthropic claiming up to 2.5 times the speed.
Knowledge cutoff is June 2026.
The model ID is claude-opus-5-5, and it is live on Anthropic's API plus Amazon Bedrock, Google Cloud, and Microsoft’s cloud stack.
Benchmarks are the usual mix of company-run tables and a few third-party suites.
On Terminal-Bench 4.0, Anthropic reports Opus 5.5 at 66.4%, ahead of OpenAI's GPT-6 Astra at 57.9% and well ahead of GPT-5.6 Sol. On FrontierCode v1.1 it edges Astra, 54.4% to 53.3%. On CursorBench 4.0 it scores 57.8% against Sol's 41.7%. On GDPval-AA v2.1, a pairwise knowledge-work eval across dozens of occupations, it posts an Elo of 1846, above Fable 5.1 at 1735 and Opus 5 at 1708. Humanity’s Last Exam with tools is listed at 67.7% versus Astra’s 57.2%. The gaps are not uniform. Astra still leads Terminal-Bench-Science 0.1, 64.6% to 58.7% , and it edges Opus 5.5 on AutomationBench. Anthropic itself notes that margins at this level are less reliable than the charts imply, and that production safeguards can zero out some cyber and biology tasks or hand them to older fallback models.
The product story underneath the leaderboard is narrower.
Anthropic is selling long-running coding agents, computer use, and professional knowledge work.
Early tester quotes in the launch post describe fewer steps, fewer tokens, and shorter calendars on migrations and reviews. One anecdote has a tester finishing a 680,000-line code migration in less than a day.
Another has the model cutting page load times on a web app in 39 of 40 trials, where Opus 5 made smaller changes that also altered behavior.
Those are selected stories. They match the direction of the evals more than they prove a new category of software.
Communication got as much airtime as coding.
Opus 5 drew a steady complaint that it was hard to work with over long sessions: buried answers, extra prose, instructions that drifted.
Anthropic says 5.5 puts the important information first and follows writing rules more tightly.
On X, that landed harder than some of the benchmark charts. One engineer called the writing change the single most important improvement, saying Opus 5 had been almost unusable because of how it talked. That is user color, not a measured metric, but it explains why a point-five version can feel larger than the version number.
Safeguards follow the Fable pattern more than the older Opus pattern.
Biology and cybersecurity classifiers are stricter.
Many cyber tasks get rerouted to Opus 4.8 unless a user is in Anthropic’s Cyber Verification Program. Life sciences access is gated through a separate verification track.
Thinking cannot be turned off, which Anthropic ties in part to anti-distillation measures for newer API accounts. The company also cites stronger prompt-injection results and zero data retention, plus EU AI Act watermarking.
That is the trade. Teams that used Opus as a relatively permissive workhorse will notice.
Subscription users got a quieter change that may matter more day to day than any Elo score.
Anthropic is raising five-hour usage limits on Pro, Max, and Team plans and issuing a rate-limit reset that subscribers can bank and spend later.
The announcement thread treated that as a postscript. Replies treated it as the headline. Token economics only help if the product lets you keep working through the afternoon.
What the launch does not settle is the pacing argument that framed it.
Shipping a cheaper, faster Opus that matches much of Fable on coding and office work is a commercial move, and a competitive one against Astra and Sol. It is also a test of whether "pacing the frontier" means fewer releases, slower capability jumps, or the same cadence with more paperwork and more external names on the system card.
Anthropic chose the third reading.
Sonnet 5.5 and Haiku 5.5 are already promised in the coming weeks. The model race did not pause. It got a new default.























































































































































































































































































































































































