Claude Sonnet 5.5 makes routing more important than model rankings
Anthropic’s faster, same-price Sonnet is worth testing now—but only as a bounded-work default with effort caps and an Opus fallback.
What changed
Anthropic released Sonnet 5.5 on September 28 at the same $2 input and $10 output per million token list price as Sonnet 5. It claims 30% faster generation and up to 30% lower cost per task, with immediate availability through the Claude API, AWS, Google Cloud and Microsoft Foundry.
The expensive setting
Independent evidence says effort settings matter more than the launch headline. CodeRabbit’s 44-PR run finished in less than half the time and at about 40% of Sonnet 5’s model cost, while Artificial Analysis found max effort extremely token-hungry and off the cost-performance frontier.
What we would ship
This is not a blind model-ID swap. Anthropic documents five breaking changes: teams that disabled thinking must move to between_tools, streamed text between tool calls can change shape, and non-default temperature, top_p or top_k values return a 400 error.
We would canary Sonnet 5.5 on bounded workloads, pin effort per route and preserve an Opus escalation path. Ship the routing change now; defer any fleet-wide migration until production evaluations confirm cost per completed task and error severity.
Claude Sonnet 5.5 is the first Sonnet release we would treat as the default candidate for bounded production work, but we would not replace every model with it. At Virtual Arc, we would shadow-run it on bug fixes, first-pass code review, support drafting and structured document work, then promote it only where completed-task cost, latency and failure rate beat the current route. We would cap routine traffic at low, medium or high effort and keep Opus 5.5 for ambiguous architecture, high-risk reviews and tasks where a missed error costs more than extra tokens. Anthropic’s headline says up to 30% lower cost per task, but Artificial Analysis found max effort consumed about 193,000 output tokens per index task and cost roughly 50% more per task than Sonnet 5. The operational lesson is not “upgrade everything”; it is “route aggressively, measure by completed task, and ban max effort from the default path until your own evaluations justify it.”