Claude Sonnet 5.5: Compare Cost per Finished Task

Sonnet 5.5 keeps Sonnet 5 token prices. Its claimed savings come from using fewer tokens. How to test that claim in your workflow.

Anthropic released Claude Sonnet 5.5 on September 28, keeping Sonnet 5’s token rates: $2 per million input tokens, $10 for output and $0.20 for cache reads. The company reports up to 30% lower cost per task because the model uses fewer tokens, plus more than 30% faster output generation. Those are vendor findings, not S5 Labs measurements. Anthropic’s announcement

For a team paying for repeated agent runs, the distinction matters. A lower bill for one attempt is useful only if the result meets the same acceptance criteria. A run that needs another attempt or substantial editing can erase the apparent saving.

Start with work you can grade

Our recommendation is to evaluate Sonnet 5.5 first on bounded tasks with observable outcomes: a bug fix with an existing reproduction, a document that must preserve supplied figures, or an extraction job with a known answer. Keep the task instructions and available tools consistent across models. Record failed attempts as part of the cost.

A practical comparison sheet can stay small:

RecordWhat it reveals
Accepted resultWhether the output meets the task’s requirements
Total model charges across attemptsWhether retries consume the initial saving
Elapsed time to acceptanceWhether faster output improves the actual workflow
Human correction timeWhether review effort has moved onto the operator

Choose the acceptance rule before reading the outputs. For code, inspect the change and run the relevant checks. For a spreadsheet or report, verify the underlying values as well as the presentation. Count a polished but incorrect artifact as a failure.

Effort belongs in the comparison

Anthropic positions Sonnet 5.5 for everyday, well-defined work and Opus 5.5 for more complex, open-ended judgment. It says lower Sonnet effort settings offer its strongest cost advantage; increasing effort can bring the cost closer to Opus. Launch analysis and effort discussion

That supports testing a modest effort setting before making a more expensive one the default. It does not establish the best setting for your application. Keep the chosen effort in your evaluation record so a later change does not look like an unexplained model regression.

Routing can then follow observed failure patterns. If a class of tasks repeatedly needs architectural decisions or extensive correction, compare a stronger model on that class. Avoid escalating every task simply because one difficult example failed; equally, do not retain a cheaper route when its review burden consistently exceeds the saving.

Check the migration before switching traffic

Anthropic lists the API identifier as claude-sonnet-5-5. Its announcement says users running with thinking disabled must move to between_tools, which leaves up-front thinking off. Review the linked migration documentation against your request format before changing production traffic. Availability and migration notes

S5 Labs has not run a comparative evaluation of this release. The useful next step is a small, repeatable trial against the work you already understand, with a retained baseline and a clear acceptance threshold. Compare the total cost of an accepted result before changing the default model.

Continue reading.

Insight3 min read

Holo4 Combines Computer Use and Tools, with Different Licenses

Holo4 spans GUI, code, MCP and API workflows. Its 27B and 35B-A3B checkpoints have different licenses. What to check before evaluation.

Insight27 min read

OpenAI DevDay 2026: Always-On Dots, GPT-6.1 Sol and Uneven Availability

OpenAI DevDay 2026 announced dots, GPT-6.1 Sol, Ultrafast, Agents API computer use and plugin extensions. What is live now, what it costs and what to test.

Insight3 min read

Perceptron Mk1.5 Adds Video Tracking for Embodied Agents

Perceptron Mk1.5 adds audio and video tracking. A September 28 assessment of its hosted API, context budget and integration constraints.