AI news

Claude Haiku 5.5 Pricing: 90% Off Haiku 4.5, Two Catches

·

Anthropic shipped Claude Haiku 5.5 on October 7, 2026, and the headline number is blunt: $0.10 per million input tokens, $0.50 per million output tokens — 90% below Haiku 4.5, and exactly level with OpenAI's GPT-6 Luna. A Claude model is now priced like a budget model, undercutting DeepSeek's flash on list price, which is a sentence nobody expected to write this year. Whether that number survives contact with your workload depends on two pieces of fine print, and both are on this page.

Short answer

  • Claude Haiku 5.5 (model ID claude-haiku-5-5) launched October 7, 2026 at $0.10/$0.50 per million tokens for prompts up to 100K tokens — 90% below Haiku 4.5's $1.00/$5.00, tied to the cent with GPT-6 Luna.
  • Catch one: prompts over 100K tokens are billed at $0.50/$2.50 — 5× the short-prompt tier. Anthropic says ~90% of historical Haiku requests stayed under 100K; if your work is long-context, this is not the cheap model.
  • Catch two: the new tokenizer bills roughly 30% more tokens for the same text than Haiku 4.5. Our math puts the effective cut at ~87% for short-prompt work and ~35% for long prompts. Anthropic's own blended claim: "around 75% less to run."
  • Same-day extras: Sonnet 5.5 cache reads halved ($0.20 → $0.10), and Claude Max/Team subscribers get $100–$500/month in API credits rolling out this week.

The price, exactly

Official list prices per million tokens, from the launch post and platform docs, with the two prompt-size tiers side by side:

Haiku 5.5 (≤100K prompt)Haiku 5.5 (>100K prompt)Haiku 4.5Sonnet 5.5
Input$0.10$0.50$1.00$2.00
Output$0.50$2.50$5.00$10.00
Cache read$0.01$0.05$0.10$0.10
Cache write (5m)$0.125$0.625$1.25$2.50

The tier is set per request by prompt length. Anthropic's own justification for the structure: prompts up to 100K tokens "make up around 90% of requests to our previous Haiku model." So the cheap tier covers the bulk of historical usage — but the model advertises a 1M-token context window, and everything past the first 100K bills at five times the rate. A 400K-token prompt does not pay $0.10 on its first 100K and $0.50 on the rest; the request prices out at the long tier. The 1M window is real, but the budget price only patrols the first 10% of it.

Then there's the tokenizer. Haiku 5.5 moves to the updated tokenizer already used by Sonnet 5.5 and Opus 5.5, and Anthropic's docs state it plainly: the same text counts as approximately 30% more tokens than on Haiku 4.5. That converts the nominal 90% cut into roughly an 87% effective cut on short-prompt input ($0.10 × 1.3 = $0.13 per old-equivalent million), and the nominal 50% cut on long prompts into roughly 35% ($0.50 × 1.3 = $0.65 vs $1.00). Anthropic's own blended claim — "on average, it now costs around 75% less to run" — folds in the same adjustment at the task level. None of this makes the discount fake; 87% is enormous. It just means the sticker says 90% and the invoice will say less.

One more lever: the Batch API takes 50% off everything, putting batch jobs at $0.05/$0.25 for short prompts — as cheap as anything in the mainstream market.

Against the budget-model field

OpenRouter's listed prices, retrieved October 8, 2026 — the marketplace where these models are actually comparison-shopped, USD per million tokens:

ModelInputOutputContext
Claude Haiku 5.5 (≤100K tier)$0.10$0.501M
OpenAI GPT-6 Luna$0.10$0.501.05M
Z.ai GLM-5.3-Flash$0.15$0.501M
Qwen3.8 Flash$0.15$0.471M
Xiaomi MiMo-V2.6-Flash (new)$0.14$0.281.05M
DeepSeek V4.1 Flash$0.30$1.201M
Google Gemini 3.8 Flash$0.75$3.751M

Read it in two halves. For short prompts, Haiku 5.5 ties GPT-6 Luna as the cheapest input price in the mainstream tier, and undercuts DeepSeek V4.1 Flash by 3× and Gemini 3.8 Flash by 7.5×. On output it is tied with Luna and GLM-5.3-Flash, and it is not the absolute floor: the newly released open-weights MiMo-V2.6-Flash bills $0.28, and Qwen3.8 Flash $0.47. Below this tier sit niche open-weights listings like Ling 3.0 Flash ($0.021/$0.063) that trade provenience and context size for price — a different buying decision.

Cross the 100K line and the table flips. At the long tier, GLM-5.3-Flash and Qwen3.8 Flash bill ~3× less on input and 5× less on output than Haiku 5.5. So "cheapest small model" now has a conditional answer: for short-prompt volume work, Anthropic is at the front of the race; for long-context budget work, the Chinese budget tier keeps the crown it already held — our GLM-5.3-Flash vs DeepSeek V4.1 Flash comparison still stands as the reference for that segment. Buying direct can also beat marketplace prices: Zhipu's official CNY pricing is covered in our GLM-5.3-Flash API price breakdown, and DeepSeek's off-peak discounts in the off-peak pricing guide.

And for the workload Haiku is explicitly built for — classification, extraction, routing over a stable system prompt — the number that matters is cache reads at $0.01 per million. Once your prompt prefix is warm, repeated high-volume calls are effectively an order of magnitude under the headline input price.

What $0.10 buys

Positioning, per the launch post: high-volume, latency-sensitive tasks — summaries, compactions, database queries, classification — plus subagent work alongside Opus 5.5 and Sonnet 5.5, and speed-sensitive uses like live support and browser use. Anthropic calls it their fastest model to date (at standard speed; Opus models in Fast Mode still run faster). It's the first Haiku-class model with an adjustable effort setting, ships a 1M context with 128K max output (300K via the Batch API beta), carries a June 2026 knowledge cutoff, and is guaranteed available no sooner than October 7, 2027 — a one-year runway for production adoption.

The published benchmarks, all vendor-run, versus same-price GPT-6 Luna:

BenchmarkHaiku 5.5GPT-6 LunaHaiku 4.5
GDPval-AA v2.1 (44-occupation agent eval)16201437735
OSWorld 2.1 (offline subset)72.4%48.9%15.7%
Humanity's Last Exam (no tools)45.9%—10.2%
Terminal-Bench 4.039.2%16.4%0.0%

Confidence: medium-high — these are Anthropic's own numbers from the launch post, not independent runs. At the same list price as Luna, every chart Anthropic published favors Haiku. The honest boundary is in the same post: Sonnet 5.5 scores 70.6% on Terminal-Bench and remains the right tool for complex agentic coding — Haiku 5.5 is for narrowly scoped, high-volume steps where the previous options were cost-prohibitive. Anthropic's customer color points the same direction: AlphaSense reported its 8M-calls-per-week document-QA product improving from 0.76 to 0.84 versus Haiku 4.5, and Asana cited ~30% lower task latency. Curated customer quotes — direction, not gospel.

The other two announcements inside the announcement

Two line items buried under the Haiku launch may matter more to existing Claude users than the model itself.

Sonnet 5.5 cache reads are halved — $0.20 to $0.10 per million, effective immediately. Because cache reads make up a large share of token consumption in agent loops, Anthropic puts the practical effect at around 20% off most agentic workloads on Sonnet 5.5. If you run Claude agents, this is the line item to re-check your bill against; the three places heavy users still overpay are mapped in our Claude API heavy-user cost guide, and two of them just moved.

Max and Team subscribers get monthly API credits — $100 for Max 5x, $200 for Max 20x, up to $500 pooled for Team, rolling out this week. The credits work on any Claude Platform model via API, Managed Agents, the Agent SDK, and the playground; they do not cover Claude Code or app overage, they don't apply on Bedrock, Google Cloud, or Microsoft Foundry, and unused balance expires at the end of each billing cycle. The practical read: a Max subscriber can now run Haiku 5.5 experiments for free — $100 at $0.10/M input is a billion input tokens of headroom. The timing is pointed, coming the same week Google cuts its free-tier model list to Flash-Lite only (October 9): the free-API landscape is being redrawn in both directions at once. Which doors are still open is tracked in our free LLM API tiers guide.

Who should switch today

  • High-volume, short-prompt pipelines on Haiku 4.5 (classification, extraction, compaction, routing): yes. The effective cut is ~87%, migration is a model-ID change, and the one-year availability guarantee de-risks it.
  • Pipelines on GLM-5.3-Flash or Qwen3.8 Flash under 100K prompts: the price gap is gone, so the decision is capability, not cost. Run your own evals before believing vendor charts — both directions are plausible.
  • Long-context budget work (>100K prompts): no. GLM-5.3-Flash, Qwen3.8 Flash, and DeepSeek V4.1 Flash all stay materially cheaper at that tier; Haiku's long-prompt price is a convenience price, not a budget price.
  • Claude Max/Team subscribers: claim the credits and run the experiment this cycle — it's the cheapest way to answer whether the benchmark jump survives your data.

Availability is across the Claude API (claude-haiku-5-5), Amazon Bedrock (anthropic.claude-haiku-5-5), Google Cloud, and Microsoft Foundry. For per-token economics across all vendors, the cheapest LLM API provider comparison remains the reference page — as of today, the answer to "who's cheapest" just got a new contestant.

Dated caveat: all Anthropic prices, dates, model specs, and benchmark figures checked October 8, 2026 against Anthropic's launch post (anthropic.com/claude-haiku-5-5) and the Claude platform docs (platform.claude.com), fetched the same day. Competitor prices are OpenRouter listed prices retrieved October 8, 2026, and can change without notice; buying direct from each vendor may differ. Benchmarks are vendor-published, not independently verified. The effective-savings math (~87% / ~35%) is ours, derived from two official numbers — the nominal 90%/50% cuts and the ~30% token-count increase — and is labeled as such. Anthropic's "~75% average" is their claim, not ours.

Sources: Anthropic — Introducing Claude Haiku 5.5 (Oct 7, 2026) · Claude Platform Docs — Claude Haiku 5.5 overview · Claude Platform Docs — API credits for subscribers · OpenRouter — Claude Haiku 5.5 listing