AI coding tools

GLM-5.3-Flash API Price Check: Who's Still Cheap After the Launch Discount?

·

Short answer

Z.ai's own list price for GLM-5.3-Flash is $0.15/M input + $0.50/M output. Today the cheapest way to call the same model through a normal API is DeepInfra at $0.075/$0.25 — but only while its 50% discount lasts, and DeepInfra publishes no end date for it. The cheapest price not sitting on a discount timer is OpenRouter at $0.09/$0.30. Z.ai direct and Novita both charge full list.

GLM-5.3-Flash launched on August 26 with launch-window pricing, and most coverage from that week still quotes $0.075/$0.25. Three weeks later that number is mostly wrong. We checked all four hosts we could reach directly on September 17, 2026 — two of them have already rolled back to list price.

What each host charges today

All prices per 1M tokens, verified on each provider's own page or API on Sept 17:

Host Input Cached input (read) Output The catch
Z.ai (official list) $0.15 $0.03 $0.50 Cached-input storage is free for now ("limited-time")
Novita $0.15 $0.03 $0.50 Launch discount is gone
OpenRouter $0.09 $0.018 $0.30 Batch endpoint: $0.075/$0.25
DeepInfra $0.075 $0.015 $0.25 List is $0.15/$0.50 — you're buying a 50% discount with no stated expiry

Specs are identical wherever you go, since it's the same model: 320B total parameters, 18B activated, a 1M-token context window, native multimodal input — all from Z.ai's own model documentation. OpenRouter's listing shows a slightly larger 1.31M context on its standard endpoint; Z.ai's docs say 1M, so plan around 1M.

Sources: Z.ai pricing page and GLM-5.3-Flash model doc; Novita's model page; OpenRouter's models API; DeepInfra's model page.

Why the same model has three different price tags

Z.ai set list at $0.15/$0.50 and shipped with roughly half-off launch pricing. That window has closed — Z.ai's pricing table now shows only $0.15/$0.50, and Novita's card shows the same numbers where $0.075/$0.25 used to sit.

DeepInfra is the outlier: its page openly shows the strikethrough list price next to "50% off" at $0.075/$0.25, with the discount flag set in its model data but discount_ends_at: null. That's a real, currently-billable price — just not a durable one. If you build a budget on it, re-check it like you'd check milk.

One more data point we could not verify: aggregator Artificial Analysis lists a Bitdeer endpoint at $0.05/M input, which would undercut everything above. We couldn't reach Bitdeer's pricing directly today, so treat that as unconfirmed rather than plan around it.

Every number on this page was checked Sept 17, 2026, and the two cheapest rows are the most likely to move. Discount pricing without a published end date can disappear overnight — re-verify before you commit a workload.

The discount that expires soonest is Z.ai's own — Sept 20

If you use GLM-5.3-Flash through a coding agent rather than raw API calls, Z.ai's GLM Coding Plan is running a usage campaign: from September 3 to September 20, every night between 23:00 and 09:00 Singapore time, GLM-5.3-Flash usage through ZCode consumes zero quota, and other supported agents get their quota doubled. The standing plan rule also halves point consumption during off-peak hours and all day on weekends.

That's the same shape as DeepSeek's off-peak discount we broke down earlier — shift your runs into the cheap window instead of chasing a lower list price. But this one has a hard stop: the zero-quota window ends September 20, so if you're reading this after that date, assume it's gone and re-check the campaign page.

Source: Z.ai's GLM-5.3-Flash usage campaign notice.

If even $0.15/M is more than you want to spend

Z.ai keeps two genuinely free GLMs on its pricing table: GLM-4.7-Flash and GLM-4.5-Flash are listed at $0 across input, cache, and output. They're older and weaker, but for many glue jobs — classification, summarization, simple extraction — free beats cheap. One tier down from 5.3-Flash, GLM-4.7-FlashX runs $0.07 in / $0.40 out.

For a second opinion on the budget tier of a different model family, our DeepSeek V4 cost calculator and off-peak pricing breakdown cover the same spend-reduction logic applied to DeepSeek's lineup — the off-peak window there is structural, not a campaign. And if you'd rather not pay anything at all, our guide to free LLM APIs that actually stay free filters out the one-time-trial bait.

Method note: prices were pulled from each host's own pricing surface — Z.ai's official docs, Novita's rendered model page, OpenRouter's public models API, and DeepInfra's model page data — not from third-party roundups. Where a host showed both list and discounted prices (DeepInfra), we quote the currently-billable price and flag the discount separately.

Checked September 17, 2026. Prices are what the four hosts' own pages showed that day; we don't track them afterward.