8 October 2026

Claude Haiku 5.5 pricing makes context limits a production concern

Anthropic’s new small model is cheap enough to reroute real workloads, but its fivefold price jump above 100,000 prompt tokens makes context control an architectural requirement.

Virtual Arc · Editorial image

What launched

Anthropic released Claude Haiku 5.5 on October 7 with a 1M-token context window, up to 128K output tokens and adjustable effort. It is generally available through the Claude API, Amazon Bedrock, Google Cloud and Microsoft Foundry. In Anthropic’s own evaluation table, Haiku 5.5 scored 39.2% on Terminal-Bench 4.0 and 72.4% on the OSWorld 2.1 offline subset; Sonnet 5.5 remains materially stronger on the harder coding benchmark.

The hidden cliff

List pricing starts at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. Above that boundary, both rates rise fivefold, and Anthropic says the new tokenizer counts the same text as roughly 30% more tokens than Haiku 4.5.

Our deployment plan

We would trial Haiku 5.5 as a routed worker for classification, extraction, compaction, summaries and narrow subagents—not as the lead model for long, ambiguous workflows. The promotion gate is lower cost per successful task at acceptable quality and p95 latency on our own production traces.

Our take

Claude Haiku 5.5 is the rare small-model launch that should change production routing now, not next quarter. Virtual Arc would not replace a lead reasoning or coding model with it; we would put it behind a router for bounded, high-volume work such as classification, extraction, compaction, summaries and narrow subagents, then judge it on cost per successful task and p95 completion time. The advertised $0.10/$0.50 per million input/output tokens applies only while prompts stay at or below 100,000 tokens; above that line both rates jump fivefold, and Anthropic says the newer tokenizer counts the same text as roughly 30% more tokens than Haiku 4.5. That makes context discipline part of the architecture, not a billing footnote. We would run a one-week shadow evaluation on real traces, cap context and effort, and promote Haiku only where quality holds. For long, messy agent loops, we would keep the stronger model and avoid turning a cheap worker into an expensive generalist.

Sources
  1. Introducing Claude Haiku 5.5
  2. Claude Haiku 5.5 – Claude Platform Docs
  3. Claude Haiku 5.5 – Amazon Bedrock
  4. Claude Haiku 5.5 on Google Cloud

← All posts