AI news

Claude Haiku 5.5: Four Ways to Actually Pay Less

·

Haiku 5.5 is already the cheapest Claude model ever listed — $0.10 per million input tokens at the short-prompt tier. But the list price is still not the bill, and for this model specifically, Anthropic shipped four separate ways to pay less on the same day it launched: subscriber credits that function as a monthly grant, a 50% batch discount, a cache-read price that undercuts its own input rate by 10×, and a pricing cliff you can plan around. This page is the action list — what each lever requires, what it saves, and where it doesn't apply.

Short answer

  • Claude Max and Team subscribers get free monthly API credits: $100/mo on Max 5x, $200/mo on Max 20x, up to $500/mo pooled on Team. They cover the Claude API, Managed Agents, the Agent SDK, and the playground — they do not cover Claude Code or extra usage in the Claude apps, and unused credits expire at the end of each billing cycle. Link a Claude Console organization to your plan to claim them.
  • Batch API cuts the bill 50%: Haiku 5.5 batch pricing is $0.05/$0.25 per million tokens (≤100K tier) — the cheapest mainstream batch tier we've tracked, level with GPT-6 Luna's batch rate.
  • Cache reads cost $0.01/MTok (≤100K tier) — one-tenth of the input price. Workloads with a stable system prompt effectively run at the cache rate after warm-up.
  • The 100K-token cliff is avoidable: prompts over 100K tokens bill at $0.50/$2.50 — 5×. Splitting long documents into sub-100K chunks, or letting Haiku do the compaction itself, keeps you on the cheap tier.

The credits: as close to free as Claude gets

The credits program rolled out the same week as Haiku 5.5, and it changes the math for anyone already paying for Claude Max or Team. The mechanics, from Anthropic's credits documentation: you link a Claude Console organization to your Max or Team plan, credits arrive each billing cycle (monthly for annual plans), and you can spend them on the Claude API, Claude Managed Agents, the Claude Agent SDK, and the playground. No payment method is required on Claude Platform to use them.

Three exclusions matter. First, Claude Code is not covered — if your plan is really a Claude Code subscription, this grant doesn't touch it. Second, extra usage in the Claude apps isn't covered either. Third, the credits only apply on the Claude Platform itself: they can't be spent on Amazon Bedrock, Google Cloud, or Microsoft Foundry deployments. And the credits are use-it-or-lose-it — they expire at the end of each billing cycle rather than rolling over.

What $100 a month buys at Haiku 5.5 short-tier rates: a billion input tokens, or two hundred million output tokens. Even mixing in the realistic ratio of 3:1 input to output, that is on the order of 400 million tokens of monthly throughput covered by a plan many subscribers already hold. For a classification, extraction, or routing workload, that is rarely pocket change.

Batch and cache: the two structural discounts

The Batch API takes a flat 50% off input and output on Haiku 5.5: $0.05/$0.25 per million tokens at the ≤100K tier, $0.25/$1.25 above it. If your workload tolerates a 24-hour window — nightly digests, bulk classification, embedding-style sweeps — batch is the first lever, and it stacks with nothing else needed.

Prompt caching is the second, and for high-volume repeated-context work it's the bigger one. Cache reads on Haiku 5.5 cost $0.01 per million tokens at the ≤100K tier ($0.05 above it), against $0.10/$0.50 for uncached input. Writes cost $0.125 (5-minute TTL) or $0.20 (1-hour) per million. The economics only work if your prompt prefix is genuinely stable — the same long system prompt or retrieved context reused across many requests. If your router or middleware rewrites prompts per request, you're paying write prices for reads that never hit. We covered that failure mode, and the router trap that causes it, in the heavy-user cost breakdown.

Planning around the 100K cliff

The tiered pricing is the one design choice that punishes careless long-context use: a 120K-token prompt bills its entire input at $0.50 per million, not just the portion above 100K. The fix is architectural, and conveniently, Haiku 5.5 is officially positioned for it — summarization and compaction are among its stated use cases. Split the corpus, summarize each chunk with Haiku, and reason over the summaries: the sub-100K tier prices the whole pipeline. For workloads that genuinely need full 1M-token context in one request, the long tier is the honest cost of admission — and at that point, Chinese budget models are materially cheaper for the same tokens. That comparison, with current prices, is in our launch-day price breakdown.

When to not bother

Two cases where the levers don't change the answer. If your volume is low and bursty — a few thousand requests a month — the engineering time to implement batching and cache discipline will cost more than the tokens save; the subscriber credits alone probably cover you. And if you're not on Claude at all: GLM-5.3-Flash and Qwen3.8 Flash remain cheaper on output at list price, DeepSeek's off-peak windows discount its own flash model further, and the standing cheapest-provider comparison is the honest starting point. Haiku 5.5's case is capability at a budget price, not the budget price itself. Which free doors are open across providers right now — including Google's free-tier contraction on October 9 — is tracked in our free-tier guide, and Google's specific cut is covered here.

Caveat: Written October 8, 2026, against three Anthropic sources fetched the same day — the Haiku 5.5 launch announcement, the Claude Platform docs model overview, and the API-credits-for-subscribers documentation. The credit amounts ($100/$200/$500), coverage rules, expiry behavior, batch rates, and cache rates all come from those pages. Program terms can change without notice; the credits page is the authoritative reference. The "400 million tokens" throughput estimate is our arithmetic from the listed rates, labeled as such.

Sources: Claude Platform Docs — API credits for subscribers · Anthropic — Introducing Claude Haiku 5.5 (Oct 7, 2026) · Claude Platform Docs — Claude Haiku 5.5 overview