Claude Haiku 5.5: A 90% List-Price Cut, a 75% Real One, and Five Ways to Get a 400

Haiku 5.5 lists at a tenth of Haiku 4.5's price; Anthropic says about 75% less in practice. The gap, the 400s after switching, and Sonnet 5.5 vs Luna.

Anthropic released Claude Haiku 5.5 on October 7 at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, against $1 and $5 for Haiku 4.5. The same announcement says the new model “on average” costs “around 75% less to run”. Both numbers are correct, and the gap between them is worth working through before changing a model ID. Anthropic’s announcement

S5 Labs has not run Haiku 5.5. Every benchmark and speed figure here is Anthropic’s, and Anthropic chose which competitor to put in its tables.

Two prices, one tokenizer

The list prices are a straight tenth of Haiku 4.5’s for prompts up to 100,000 tokens. Above that threshold the rates are five times higher, which still puts them at half of Haiku 4.5’s. Pricing

Per 1M tokensHaiku 4.5Haiku 5.5, prompt ≤100kHaiku 5.5, prompt >100k
Input$1.00$0.10$0.50
Output$5.00$0.50$2.50
Cache read$0.10$0.01$0.05
5-minute cache write$1.25$0.125$0.625
1-hour cache write$2.00$0.20$1.00
Batch input / output$0.50 / $2.50$0.05 / $0.25$0.25 / $1.25

Anthropic’s footnote explains how 90% off becomes 75% off. On Haiku 4.5, it says, 90% of requests were under 100,000 tokens. Weight the two price cuts by that mix, counting requests rather than tokens, and the list-price saving is about 86%. The footnote then says the 75% figure “also accounts for changes between Haiku 4.5 and Haiku 5.5 in how many tokens are used to complete a given piece of work”, because Haiku 5.5 “has an updated tokenizer” and “uses slightly more tokens per task”. Anthropic’s announcement

The docs are more specific than the footnote. Haiku 5.5 uses the tokenizer introduced with Claude Opus 4.7, and the migration guide states that “the same input text produces approximately 30% more tokens on Claude Haiku 5.5 than on Claude Haiku 4.5”, with the exact increase depending on content. Migration guide The models overview puts it another way: a million tokens on the new tokenizer holds roughly 555,000 English words, where the old one held about 750,000. Models overview

A 30% token increase does not close the gap between 86% and 75% on its own. A 75% saving means the bill is a quarter of what it was; an 86% list-price cut means the rate is a seventh. For those to reconcile, Haiku 5.5 would have to spend something like 1.8 times the tokens Haiku 4.5 spent on the same work, if the request mix held. That is our arithmetic, not Anthropic’s, and the footnote does not break its figure down. The likely second contributor is adaptive thinking, which is on by default and counts as output: Haiku 4.5 ran without thinking unless asked, and Haiku 5.5 thinks at medium effort unless told otherwise. Effort Anthropic’s prompting guide gives a sense of the scale: in its tests of a long agent prompt, moving from low to medium effort “more than doubled the output tokens for each attempt”. Prompting Claude Haiku 5.5

So the effective price depends on a setting that did not exist on the old model. Haiku 5.5 is the first Haiku with the effort parameter: five levels from low to max, medium the default. Thinking can still be turned off with thinking: {"type": "disabled"} at high effort or below; at xhigh and max that request returns a 400. Effort For classification and routing, Anthropic says low is “the cheapest and fastest level” and suits “simple, high-volume requests”. Prompting Claude Haiku 5.5 One small item goes the other way: the hidden system prompt that tool declarations add is 286 tokens on Haiku 5.5 against 496 on Haiku 4.5. Pricing

A worked example: one million triage calls

Take a support-ticket classifier. The assumptions: a 2,000-token system prompt on Haiku 4.5’s tokenizer, served from the five-minute cache; a 500-token ticket; a 150-token structured answer. For Haiku 5.5, apply the 30% tokenizer increase across the board (2,600, 650 and 195 tokens), and allow 300 thinking tokens per call at medium effort, which is a guess, since Anthropic publishes no per-call thinking figure. Cache writes, tool definitions and retries are left out. All prompts stay under 100,000 tokens.

Per 1M callsCache readFresh inputOutput incl. thinkingTotal
Haiku 4.5$200$500$750$1,450
Haiku 5.5, medium, 300 thinking tokens$26$65$248$339
Haiku 5.5, thinking disabled$26$65$98$189
Sonnet 5.5, same tokens, cache read at $0.10$260$1,300$4,950$6,510
GPT-6 Luna, same token counts$26$65$248$339

On those assumptions Haiku 5.5 comes in 77% below Haiku 4.5 with thinking on, close to Anthropic’s average, and 87% below with thinking off. The thinking allowance is the swing factor: every 100 thinking tokens per call adds $50 per million calls, and a classifier that reasons for 1,000 tokens before answering would cut the saving from 77% to about 52%. Measure it with the token-counting endpoint and model set to claude-haiku-5-5, as the migration guide says, rather than scaling old counts. Migration guide

The Sonnet row uses the cache-read price Anthropic cut on the same day, from $0.20 to $0.10 per million tokens, which the announcement says makes Sonnet 5.5 “around 20% cheaper” on most agentic work. Release notes Anthropic’s announcement Our Sonnet 5.5 cost-per-task piece was written at the old rate; its method, grading accepted results and counting retries, applies here too. The Luna row assumes identical token counts, which will not hold: OpenAI’s tokenizer differs, and Luna also defaults to medium reasoning. GPT-6 Luna model page

The 100,000-token step

Haiku 5.5 is the only current Claude model priced by prompt length. Opus 5.5 and Sonnet 5.5 charge one rate across the full million-token window; Haiku 5.5 charges five times the rate on any request whose prompt exceeds 100,000 tokens, and the higher rate covers that request’s output as well as its input and cache, even though only the prompt is measured. Pricing

That makes the threshold a cliff: a 100,000-token prompt costs $0.01 in input; a 120,000-token prompt costs $0.06. Counted on the new tokenizer, 100,000 tokens is roughly 55,000 words, and a conversation that was comfortably under the line on Haiku 4.5 can cross it after the 30% recount. GPT-6 Luna’s long-context premium applies above 272,000 input tokens, doubling input rates and raising output by half for the whole request, so between 100,000 and 272,000 tokens Luna’s input rate is a fifth of Haiku 5.5’s. GPT-6 Luna model page For long-document extraction, that is the comparison to run before the benchmark tables.

What returns a 400 after the switch

The model ID is claude-haiku-5-5 on the Claude API, Google Cloud, Microsoft Foundry and Claude Platform on AWS, and anthropic.claude-haiku-5-5 on Amazon Bedrock. It is a fixed ID with no date suffix and no alias. Migration guide Swapping the ID in is the easy part: Anthropic’s what’s-new page lists five breaking changes, and each produces a 400 error rather than a degraded answer, with one account-age exception noted in the table. What’s new in Claude Haiku 5.5

Request on Haiku 4.5On Haiku 5.5Replacement
thinking: {"type": "enabled", "budget_tokens": N}400{"type": "adaptive"} plus output_config.effort
temperature or top_p at a non-default value, or any top_k400Omit all three; prompt for the behaviour instead
A final assistant turn as prefill400, even with thinking offEnd on a user turn; use structured outputs or enum tools for format
computer_20250124 on the Claude API or Google Cloud400computer_toolset_20260801, with the old beta header dropped
Thinking block sent back after system, tools or earlier messages changed400 on accounts created from August 31, 2026; older accounts only if they set prefix_mismatch_behaviorKeep conversations append-only

Two rows cover more code than they appear to. The sampling rule is strict: temperature must be 1 if sent, top_p must be 0.99, a top_p of 1 fails, and sending both fails. Migration guide Anthropic’s Python SDK from v1.0 removes the three parameters entirely, so an old call raises a TypeError before it reaches the network. Model deprecations Prefill was the standard way to force a JSON opening brace or skip a preamble on Haiku-class models, and it is gone with no toggle; on Bedrock, where structured outputs are not supported, the guide says to use tools instead. Migration guide Anthropic’s docs list further 400s in narrower setups: a fine-grained-tool-streaming-2025-05-14 header alongside a toolset entry, a per-message effort change to a different level while thinking is disabled, and thinking: {"type": "disabled"} at xhigh or max. Migration guide Effort

What changes without an error

The remaining changes return no error, so a test suite will not catch them.

  • Responses can start with a thinking block. Code that reads content[0].text gets the wrong block. Select by type. The thinking block arrives with an empty thinking field and a signature; to see summarized thinking, set thinking.display to "summarized".
  • Thinking counts against max_tokens. A limit sized for a Haiku 4.5 classifier can end the response after the thinking block with stop_reason: "max_tokens" and no text. Raise the limit or lower effort.
  • Refusals are new. Haiku 5.5 runs safety classifiers that return stop_reason: "refusal" with a category (cyber, frontier_llm, bio or general_harms), and the docs say benign work “can also trigger” two of them. There is no server-side fallback, and a retry “usually returns another refusal”. Pipelines that never saw this stop reason on Haiku 4.5 need a branch for it. Prompting Claude Haiku 5.5
  • Thinking blocks are account-bound. Replayed through a different account, they are dropped silently and the request succeeds without the reasoning. This matters for a multi-tenant service that stores conversations centrally.
  • Priority Tier is not supported. A capacity commitment on Haiku 4.5 does not carry over. Migration guide
  • Context and output grew. 1M tokens of context and 128k of output, from 200k and 64k on Haiku 4.5. The browser-use toolset is available on the Claude API and Google Cloud; it was not on Haiku 4.5. What’s new in Claude Haiku 5.5

The prompting guide says existing Haiku 4.5 prompts “should perform well without changes”, then documents where they do not: skipped tool calls with thinking off and JSON output requested, early stopping in long agent prompts at low effort, and code changes reported done without a check. Each comes with suggested system-prompt text. Prompting Claude Haiku 5.5

Haiku 4.5’s clock

No deprecation has been announced for Haiku 4.5. The deprecations page lists claude-haiku-4-5-20251001 as Active with a tentative retirement “not sooner than October 15, 2026”, eight days after the Haiku 5.5 launch; claude-haiku-5-5 carries a commitment of not sooner than October 7, 2027. Anthropic promises at least 60 days’ notice before retiring a publicly released model, so even a notice issued today would keep Haiku 4.5 running into December on Anthropic-operated platforms; Bedrock and Google Cloud set their own dates. Model deprecations After October 15 the window is open-ended but no longer guaranteed, and the last two Haiku retirements, 3.5 and 3, each came two months after their notices. Model deprecations

Haiku 5.5, Sonnet 5.5 or GPT-6 Luna

Up to 100,000 tokens, Haiku 5.5 and Luna have the same list price on every line: $0.10 input, $0.01 cached, $0.125 cache write, $0.50 output. GPT-6 Luna model page Pricing Our GPT-6 Sol and Luna article has Luna’s API constraints. Anthropic’s own comparison table, excerpted below, puts Haiku 5.5 ahead of Luna on every row where it lists a Luna score:

Anthropic-reportedHaiku 5.5Haiku 4.5GPT-6 LunaSonnet 5.5
OSWorld 2.1 (offline subset)72.4%15.7%48.9%83.9%
Terminal-Bench 4.039.2%0.0%16.4%70.6%
GDPval-AA v2.1 (Elo)162073514371840
FrontierCode 1.1 (Main)46.4%—42.4%52.1% (xhigh)

Anthropic’s announcement The Haiku 4.5 column says more than the Luna one: 15.7% to 72.4% on OSWorld and zero to 39.2% on Terminal-Bench describe a different class of model rather than a price revision, and teams who gave up on Haiku for agent work last year should re-test within the effort levels their budget allows. Anthropic says it is the company’s “fastest model to date at each model’s standard speed”, slower only than Opus in Fast Mode; the announcement gives no tokens-per-second figure. Our benchmarks guide covers what a vendor-selected table can and cannot tell you.

Our reading of the three options:

  • Haiku 5.5 for classification, extraction, routing and subagents where prompts stay under 100,000 tokens. Start at low effort for simple requests and medium for anything with tools, and log thinking tokens so the effective price is known rather than assumed.
  • Sonnet 5.5 where the Terminal-Bench gap matters: agentic coding and multi-step tool use that fails expensively. At $2 and $10 it is twenty times Haiku’s list price at short prompts and has to earn that on acceptance rate; the new $0.10 cache read narrows the gap on cache-heavy loops. Anthropic itself says to compare Haiku 5.5 at xhigh or max against Sonnet 5.5 “on performance, cost, and speed”, which implies the top of Haiku’s range overlaps the bottom of Sonnet’s. Effort
  • GPT-6 Luna where prompts routinely run between 100,000 and 272,000 tokens, or where the stack is already OpenAI’s and the 400s above are a migration cost in their own right. On Anthropic’s numbers it trails Haiku 5.5 at equal price; we have not found an equivalent table from OpenAI.

Opus 5.5 set part of the pattern in September, with prefix-checked thinking blocks; Anthropic’s docs tie account-bound thinking blocks to Sonnet 5.5 and Haiku 5.5. Preserved thinking Our Opus 5.5 piece has the conversation-handling detail. For a team on Haiku 4.5, the order is: recount tokens on the new model, strip sampling parameters and prefill, set effort explicitly, add a refusal branch, and run the old evaluation set at low and medium before reading the price list again.

Key details

ItemDetail
ReleasedOctober 7, 2026
Model IDclaude-haiku-5-5 (fixed, no date suffix, no alias); anthropic.claude-haiku-5-5 on Bedrock
PlatformsClaude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS
Price, prompt ≤100k$0.10 in / $0.50 out / $0.01 cache read / $0.125 5m write per 1M tokens; batch 50% off
Price, prompt >100k$0.50 in / $2.50 out / $0.05 cache read / $0.625 5m write
Anthropic’s average saving“around 75% less to run” than Haiku 4.5, after tokenizer and token-use changes
TokenizerSame as Claude 4.7 and later; “approximately 30% more tokens” for the same text than Haiku 4.5
Context / output1M tokens / 128k (300k on the Batch API with a beta header)
ThinkingAdaptive, on by default; effort levels low to max, default medium; can be disabled at high or below
Breaking changesbudget_tokens, non-default sampling parameters, assistant prefill, computer_20250124, edited history with thinking blocks
Not supportedPriority Tier; server-side fallback on refusals
Haiku 4.5 statusActive; retirement not sooner than October 15, 2026; no deprecation notice as of October 7
Same-day changeSonnet 5.5 cache reads cut from $0.20 to $0.10 per 1M tokens

Sources

Continue reading.

Insight14 min read

Decision Models Go Open Weight: Cloudflare Clef, Liquid d1 and TypeSafe Jev Compared

Cloudflare's Clef and Liquid's d1-3B are open-weight decision models. Compared with Jev and OpenAI's Decisions API on license, price, context and limits.

Insight13 min read

OpenAI Decisions API: $0.10 per Million Input Tokens and Nothing for the Answer

OpenAI's Decisions API: public beta, GPT-6 Luna only, $0.10 per million input tokens, no output charge. Where it beats Responses with structured output.

Insight3 min read

Claude Sonnet 5.5: Compare Cost per Finished Task

Sonnet 5.5 keeps Sonnet 5 token prices. Its claimed savings come from using fewer tokens. How to test that claim in your workflow.