Guides & answers

GLM Coding Plan: Worth It? Full Breakdown

·

Short answer

  • Z.ai's GLM Coding Plan costs $18/mo (Lite) on monthly billing — $12.60/mo billed yearly — and for anyone running coding agents daily, the math beats pay-as-you-go API by a wide margin.
  • The Lite plan includes 10,000 credits per week; a typical GLM-5.3-Flash agent session (400K input, 100K of them cache hits, 50K output) burns about 115 credits — that's ~85 such sessions a week on the cheapest plan.
  • Through October 7, 2026, every day from 23:00-09:00 (UTC+8), GLM-5.3-Flash via ZCode or AutoClaw consumes zero quota — unlimited off-peak usage for all paid plans, no activation needed.
  • Light users (a few short sessions a week) lose money on the subscription: at the official API rate of $0.15/M input, $0.50/M output, you'd need roughly 75M tokens a month (about 170 sessions like the example below) before the $12.60 Lite plan wins.

Z.ai's GLM Coding Plan is the subscription layer over the GLM-5.3 and GLM-5.3-Flash APIs, built for coding agents: Claude Code, OpenCode, ZCode, AutoClaw, Cursor, and some 20+ other tools. Its pricing is credit-based and genuinely different from the API's per-token bill, and there's a limited-time off-peak campaign running right now. Here's what the plan actually includes, checked against official sources on .

Plans and prices: the official numbers

There are three individual tiers. Monthly billing runs at list price; yearly billing takes 30% off (quarterly takes 20% off). All three tiers support the same models — GLM-5.3 and GLM-5.3-Flash — plus Vision, Web Search, Web Reader, and Zread MCP tools.

PlanMonthlyYearly (per mo)CreditsPositioning
Lite$18$12.6010,000 / weeklightweight iteration, small repos
Pro$80$566× Lite usageday-to-day dev, mid-sized repos
Max$168$117.6014× Lite usagepower users, priority at peak

Under the hood the limits are published precisely: Lite gets 2,000 credits per rolling 5 hours and 10,000 per week; Pro gets 12,000 per 5 hours and 60,000 weekly; Max gets 28,000 and 140,000. The 5-hour bucket refreshes dynamically, the weekly bucket resets every 7 days from subscription.

How credits burn — and the session that costs 115

Credit usage is not a flat per-request fee. Per the official formula: (input tokens × input multiplier + cached input × cached multiplier + output tokens × output multiplier) ÷ 10,000. For GLM-5.3-Flash the multipliers are 2.3 / 0.56 / 8 (input / cached input / output); for the full GLM-5.3, they're 6.9 / 1.7 / 24 — meaning GLM-5.3 costs roughly 3× more per token than Flash on the same plan.

Worked example, a real-ish agent session on GLM-5.3-Flash: 400,000 input tokens, 100,000 of them cache hits, 50,000 output.

(300,000 × 2.3 + 100,000 × 0.56 + 50,000 × 8) ÷ 10,000 = 114.6 credits

The formula treats each pricing type as its own bucket — cache-hit tokens pay the cached multiplier instead of the full input multiplier, not on top of it (the API price table prices cached input separately at $0.03/M, and double-charging would make caching pointless). On Lite's 10,000 weekly credits, that session-size fits about 87 times a week. The same session on GLM-5.3 would burn (300,000×6.9 + 100,000×1.7 + 50,000×24) ÷ 10,000 = 344 credits — nearly 3× faster. If you route most work to Flash, Lite goes far; if you insist on full GLM-5.3 everywhere, Pro starts making sense quickly.

Model routing gotcha: requests for older GLM-5.2/GLM-5.1 are automatically routed to GLM-5.3, and GLM-4.7 requests route to GLM-5.3-Flash. If your agent config still names old models, you're paying (or being metered) at the new model's multiplier — check which one your tool actually hits.

The off-peak campaign: free unlimited Flash, until October 7

From September 3 through October 7, 2026, all paid-plan users get an automatic nightly window: every day 23:00 to 09:00 the next day (UTC+8), GLM-5.3-Flash usage through ZCode or AutoClaw consumes zero quota — unlimited. Usage through other supported agents during the same window gets doubled quota. No activation; it applies by itself.

Watch the timezone: 23:00-09:00 UTC+8 is late afternoon to early morning in the US — 11:00-21:00 in New York, 16:00-01:00 in London. For American developers this "off-peak" window is actually prime working hours, which makes it unusually generous; for Asian developers it's the classic overnight batch window.

Subscription vs API: where the line sits

The official API price for GLM-5.3-Flash is $0.15/M input and $0.50/M output (cache hits $0.03/M; our four-provider price comparison has the hosted alternatives). Compare the 115-credit session above: on the API, the same token mix costs about $0.073 ($0.045 fresh input + $0.003 cache hits + $0.025 output). At that rate, you'd need around 170 such sessions a month — roughly 5.7 per day — for the API bill to reach Lite's $12.60 yearly price.

Usage profileAPI cost/month (Flash only)Best option
Casual: ~10 short sessions~$1-2API pay-as-you-go
Weekly hobby project~$5-10API, usually
Daily agent work, 3-5 sessions$25-40Lite ($12.60)
All-day GLM-5.3 heavy use$100+Pro ($56) or Max ($117.60)

The plan is cheapest-of-the-heavy-options by design: if you're already paying for multiple coding-tool subscriptions, Z.ai pitches one plan covering 20+ agent tools as consolidation. For how these seat-and-usage models compare across the industry, see what AI coding assistants actually cost, and for the model quality angle behind the plan, our GLM 5.3 Flash vs DeepSeek V4.1 Flash comparison.

Dated caveat: all prices, credit multipliers, and campaign windows verified against Z.ai's official docs and subscription page on . The off-peak campaign ends October 7, 2026 and quotas are subject to a fair-use policy — confirm current terms before subscribing. We are not affiliated with Z.ai and hold no referral arrangement (monetization frozen).

The community quality debate: what's fair to worry about

Search this plan and you'll sit next to a blunt community verdict: a thread on r/ZaiGLM titled "Don't buy GLM coding plans. Quality is atrocious." (we captured it on ; read it here). We price plans, we don't grade models — but we can bound the risk honestly instead of pretending the thread isn't there.

Two things are true at once. The same thread carries a direct rebuttal — "It's a workhorse. I get good results. I can get up to about 120k before needing a reset" — and that split matches what the pricing itself implies. Where the plan holds up: high-volume repetitive edits, boilerplate and spec'd refactors, prototypes, anything where you review the diff anyway — and the 23:00-09:00 UTC+8 window, where Flash is free-unlimited until October 7, so stress-testing it on your own workload costs nothing but the clock. Where it may genuinely disappoint: architecture-level calls, long-context reasoning (the "reset around 100k" habit in that thread is a real signal, not noise), and correctness-critical paths where a frontier model earns its keep. The 115-credit session above makes the asymmetry cheap to exploit: ~87 of those fit in Lite's weekly allowance, and off-peak Flash costs zero.

Context worth having: Z.ai's own release notes for the current GLM-5.3 generation claim higher coding-benchmark scores and better real-world behavior in tools like Claude Code and Cline — an official claim, not our measurement, and not a reply to that thread. Also check the generation on any price you read: a $10/month review online refers to the GLM-5.1 era (March 2026), other pages list "$18/$72/$160" caps or call $18 an "alternative" price. Our numbers above come from the official subscribe page on : Lite is $18/mo, $12.60/mo billed yearly. Dates and generations move faster than the articles about them — including this one.

Checked . Plan prices from z.ai/subscribe (monthly and yearly billing); credit allowances, multipliers, and routing rules from docs.z.ai devpack documentation; campaign details from the official GLM-5.3-Flash Usage Campaign notice. All times UTC+8 per the official notice.

Sources: Z.ai GLM Coding Plan subscription page · Z.ai devpack documentation · GLM-5.3-Flash Usage Campaign notice · r/ZaiGLM thread (community quality debate)