The Two-Tiered Agent Architecture: How Anthropic Designed 'Claude Haiku' to Be a Workhorse, Not a Brain

Most teams that have been running Claude for the last couple of weeks have not been choosing a single model. 

They have been splitting work. A larger model handles the judgment calls, and something cheaper is supposed to handle the volume around it: the summaries, the labels, the lookups, the steps that do not need the full model. 

That split only works if the cheaper model is actually good enough to take the handoff.

Claude Sonnet 5.5 arrived as the middle of the 5.5 line, about a week after Opus 5.5. Anthropic positioned it as the faster, lower-cost complement to Opus: stronger on well-scoped everyday tasks, bug fixes, and producing documents, slides, and spreadsheets, with a reported 30% or more speed gain over Sonnet 5 and, in their testing, up to 30% less cost per task because it used fewer tokens. On Terminal-Bench 4.0, an agentic coding evaluation, it scored 70.6% against Sonnet 5's 10.3%. 

The announcement also said a Haiku model for high-volume work would follow. 

It did, on October 7.

Claude Haiku 5.5 is that model. 

The model, which ID is claude-haiku-5-5, has a 1 million token context window and a 128,000 token output limit, and it is available on Claude.ai, the Claude Platform API, Claude Code, Amazon Web Services, Google Cloud, and Microsoft Foundry. 

On its website post, Anthropic describes it as built for high-volume, cost-sensitive work, and as the fastest model it has released, slower only than Opus models running in Fast Mode.

In practice the intended use is narrow and repetitive rather than open-ended. 

Summaries, compactions, classification, routing, extraction, and database-style queries are the workloads it is meant to absorb. It is also fast enough, on Anthropic's account, for live customer support and browser use. 

Computer use and browser use are available in beta through the updated Python and TypeScript SDKs. 

The other intended role is as a subagent next to Sonnet 5.5 or Opus 5.5 on coding work: the larger model keeps the plan, and Haiku pulls a figure out of a filing, triages a ticket, or runs a bounded search. 

That is a different job from asking Haiku to own a long agentic coding session. On Terminal-Bench 4.0 it scores 39.2%, against 0.0 percent for Haiku 4.5 and 70.6% for Sonnet 5.5. 

On FrontierCode 1.1 it is closer, at 46.4% against Sonnet 5.5's 52.1%. 

The gap is the reason Anthropic still points complex coding at Sonnet and Opus.

The same pattern shows up outside coding, where Haiku 5.5 scores higher against its predecessor.

Pricing is where the use case becomes concrete. 

For prompts up to 100,000 tokens, input is $0.10 per million tokens and output is $0.50, against $1 and $5 for Haiku 4.5. Above 100,000 tokens those rates rise to $0.50 and $2.50. Cache reads are $0.01 and $0.05 on the two tiers, cache writes $0.125 and $0.625. 

Anthropic says about 90% of requests to the previous Haiku were under 100,000 tokens, and that the new model is about 90 percent cheaper on that band and 50% cheaper above it. 

The average it quotes, around 75% less to run than Haiku 4.5, already factors in a newer tokenizer that counts roughly 30% more tokens for the same text. The saving is real on short, repeated calls. It shrinks if a workload sits in the long-prompt tier or if output is heavy.

Haiku 5.5 is also the first Haiku with an adjustable effort setting, so a caller can trade intelligence against cost on a given request rather than picking a different model. 

That matters for mixed queues: classification and routing can stay on a lower effort, and a harder extraction can be stepped up without leaving the cheap tier. It does not close the gap to Sonnet on long agentic runs.

The same release cut the price of cache reads on Sonnet 5.5 in half, from $0.20 to $0.10 per million tokens. 

Alignment evaluations moved in the same direction as capability. 

Anthropic reports major improvements over Haiku 4.5 on almost all of its alignment tests, with fewer instances of misaligned behavior. Cybersecurity safeguards are tighter than on Haiku 4.5 and still looser than on Sonnet 5.5; penetration-testing requests are blocked. Biology safeguards match those on Sonnet 5, Sonnet 5.5, and Opus 5.

The practical split, then, is the one Sonnet 5.5's release already implied. Sonnet 5.5 remains the model for well-scoped coding, documents, and agentic sessions, and it is now cheaper to keep in a long context because of the cache-read cut. 

Haiku 5.5 is the model for the surrounding traffic: routing, labeling, summarizing, compacting, and the bounded subtasks a larger model can hand off. 

The broader point is that Haiku 5.5 is not really competing with Sonnet 5.5 on the same terms. Performance and capability wise, it's on a different level.

Its value is in making the smaller decisions around a larger model cheap enough to run continuously. 

For teams building agentic systems, that could make the model less a destination than a workhorse: handling the thousands of routine calls that surround the comparatively few decisions where Sonnet or Opus is actually needed.

 

Published