Anthropic released Claude Sonnet 5.5 on September 28, keeping Sonnet 5’s token rates: $2 per million input tokens, $10 for output and $0.20 for cache reads. The company reports up to 30% lower cost per task because the model uses fewer tokens, plus more than 30% faster output generation. Those are vendor findings, not S5 Labs measurements. Anthropic’s announcement
For a team paying for repeated agent runs, the distinction matters. A lower bill for one attempt is useful only if the result meets the same acceptance criteria. A run that needs another attempt or substantial editing can erase the apparent saving.
Start with work you can grade
Our recommendation is to evaluate Sonnet 5.5 first on bounded tasks with observable outcomes: a bug fix with an existing reproduction, a document that must preserve supplied figures, or an extraction job with a known answer. Keep the task instructions and available tools consistent across models. Record failed attempts as part of the cost.
A practical comparison sheet can stay small:
| Record | What it reveals |
|---|---|
| Accepted result | Whether the output meets the task’s requirements |
| Total model charges across attempts | Whether retries consume the initial saving |
| Elapsed time to acceptance | Whether faster output improves the actual workflow |
| Human correction time | Whether review effort has moved onto the operator |
Choose the acceptance rule before reading the outputs. For code, inspect the change and run the relevant checks. For a spreadsheet or report, verify the underlying values as well as the presentation. Count a polished but incorrect artifact as a failure.
Effort belongs in the comparison
Anthropic positions Sonnet 5.5 for everyday, well-defined work and Opus 5.5 for more complex, open-ended judgment. It says lower Sonnet effort settings offer its strongest cost advantage; increasing effort can bring the cost closer to Opus. Launch analysis and effort discussion
That supports testing a modest effort setting before making a more expensive one the default. It does not establish the best setting for your application. Keep the chosen effort in your evaluation record so a later change does not look like an unexplained model regression.
Routing can then follow observed failure patterns. If a class of tasks repeatedly needs architectural decisions or extensive correction, compare a stronger model on that class. Avoid escalating every task simply because one difficult example failed; equally, do not retain a cheaper route when its review burden consistently exceeds the saving.
Check the migration before switching traffic
Anthropic lists the API identifier as claude-sonnet-5-5. Its announcement says users running with thinking disabled must move to between_tools, which leaves up-front thinking off. Review the linked migration documentation against your request format before changing production traffic. Availability and migration notes
S5 Labs has not run a comparative evaluation of this release. The useful next step is a small, repeatable trial against the work you already understand, with a retained baseline and a clear acceptance threshold. Compare the total cost of an accepted result before changing the default model.
