Guides & answers
GLM Coding Plan: Worth It? Full Breakdown
Short answer
- Z.ai's GLM Coding Plan costs $18/mo (Lite) on monthly billing — $12.60/mo billed yearly — and for anyone running coding agents daily, the math beats pay-as-you-go API by a wide margin.
- The Lite plan includes 10,000 credits per week; a typical GLM-5.3-Flash agent session (400K input, 100K of them cache hits, 50K output) burns about 115 credits — that's ~85 such sessions a week on the cheapest plan.
- Through October 7, 2026, every day from 23:00-09:00 (UTC+8), GLM-5.3-Flash via ZCode or AutoClaw consumes zero quota — unlimited off-peak usage for all paid plans, no activation needed.
- Light users (a few short sessions a week) lose money on the subscription: at the official API rate of $0.15/M input, $0.50/M output, you'd need roughly 75M tokens a month (about 170 sessions like the example below) before the $12.60 Lite plan wins.
Z.ai's GLM Coding Plan is the subscription layer over the GLM-5.3 and GLM-5.3-Flash APIs, built for coding agents: Claude Code, OpenCode, ZCode, AutoClaw, Cursor, and some 20+ other tools. Its pricing is credit-based and genuinely different from the API's per-token bill, and there's a limited-time off-peak campaign running right now. Here's what the plan actually includes, checked against official sources on .
Plans and prices: the official numbers
There are three individual tiers. Monthly billing runs at list price; yearly billing takes 30% off (quarterly takes 20% off). All three tiers support the same models — GLM-5.3 and GLM-5.3-Flash — plus Vision, Web Search, Web Reader, and Zread MCP tools.
| Plan | Monthly | Yearly (per mo) | Credits | Positioning |
|---|---|---|---|---|
| Lite | $18 | $12.60 | 10,000 / week | lightweight iteration, small repos |
| Pro | $80 | $56 | 6× Lite usage | day-to-day dev, mid-sized repos |
| Max | $168 | $117.60 | 14× Lite usage | power users, priority at peak |
Under the hood the limits are published precisely: Lite gets 2,000 credits per rolling 5 hours and 10,000 per week; Pro gets 12,000 per 5 hours and 60,000 weekly; Max gets 28,000 and 140,000. The 5-hour bucket refreshes dynamically, the weekly bucket resets every 7 days from subscription.
How credits burn — and the session that costs 115
Credit usage is not a flat per-request fee. Per the official formula: (input tokens × input multiplier + cached input × cached multiplier + output tokens × output multiplier) ÷ 10,000. For GLM-5.3-Flash the multipliers are 2.3 / 0.56 / 8 (input / cached input / output); for the full GLM-5.3, they're 6.9 / 1.7 / 24 — meaning GLM-5.3 costs roughly 3× more per token than Flash on the same plan.
Worked example, a real-ish agent session on GLM-5.3-Flash: 400,000 input tokens, 100,000 of them cache hits, 50,000 output.
(300,000 × 2.3 + 100,000 × 0.56 + 50,000 × 8) ÷ 10,000 = 114.6 credits
The formula treats each pricing type as its own bucket — cache-hit tokens pay the cached multiplier instead of the full input multiplier, not on top of it (the API price table prices cached input separately at $0.03/M, and double-charging would make caching pointless). On Lite's 10,000 weekly credits, that session-size fits about 87 times a week. The same session on GLM-5.3 would burn (300,000×6.9 + 100,000×1.7 + 50,000×24) ÷ 10,000 = 344 credits — nearly 3× faster. If you route most work to Flash, Lite goes far; if you insist on full GLM-5.3 everywhere, Pro starts making sense quickly.
The off-peak campaign: free unlimited Flash, until October 7
From September 3 through October 7, 2026, all paid-plan users get an automatic nightly window: every day 23:00 to 09:00 the next day (UTC+8), GLM-5.3-Flash usage through ZCode or AutoClaw consumes zero quota — unlimited. Usage through other supported agents during the same window gets doubled quota. No activation; it applies by itself.
Watch the timezone: 23:00-09:00 UTC+8 is late afternoon to early morning in the US — 11:00-21:00 in New York, 16:00-01:00 in London. For American developers this "off-peak" window is actually prime working hours, which makes it unusually generous; for Asian developers it's the classic overnight batch window.
Subscription vs API: where the line sits
The official API price for GLM-5.3-Flash is $0.15/M input and $0.50/M output (cache hits $0.03/M; our four-provider price comparison has the hosted alternatives). Compare the 115-credit session above: on the API, the same token mix costs about $0.073 ($0.045 fresh input + $0.003 cache hits + $0.025 output). At that rate, you'd need around 170 such sessions a month — roughly 5.7 per day — for the API bill to reach Lite's $12.60 yearly price.
| Usage profile | API cost/month (Flash only) | Best option |
|---|---|---|
| Casual: ~10 short sessions | ~$1-2 | API pay-as-you-go |
| Weekly hobby project | ~$5-10 | API, usually |
| Daily agent work, 3-5 sessions | $25-40 | Lite ($12.60) |
| All-day GLM-5.3 heavy use | $100+ | Pro ($56) or Max ($117.60) |
The plan is cheapest-of-the-heavy-options by design: if you're already paying for multiple coding-tool subscriptions, Z.ai pitches one plan covering 20+ agent tools as consolidation. For how these seat-and-usage models compare across the industry, see what AI coding assistants actually cost, and for the model quality angle behind the plan, our GLM 5.3 Flash vs DeepSeek V4.1 Flash comparison.
The community quality debate: what's fair to worry about
Search this plan and you'll sit next to a blunt community verdict: a thread on r/ZaiGLM titled "Don't buy GLM coding plans. Quality is atrocious." (we captured it on ; read it here). We price plans, we don't grade models — but we can bound the risk honestly instead of pretending the thread isn't there.
Two things are true at once. The same thread carries a direct rebuttal — "It's a workhorse. I get good results. I can get up to about 120k before needing a reset" — and that split matches what the pricing itself implies. Where the plan holds up: high-volume repetitive edits, boilerplate and spec'd refactors, prototypes, anything where you review the diff anyway — and the 23:00-09:00 UTC+8 window, where Flash is free-unlimited until October 7, so stress-testing it on your own workload costs nothing but the clock. Where it may genuinely disappoint: architecture-level calls, long-context reasoning (the "reset around 100k" habit in that thread is a real signal, not noise), and correctness-critical paths where a frontier model earns its keep. The 115-credit session above makes the asymmetry cheap to exploit: ~87 of those fit in Lite's weekly allowance, and off-peak Flash costs zero.
Context worth having: Z.ai's own release notes for the current GLM-5.3 generation claim higher coding-benchmark scores and better real-world behavior in tools like Claude Code and Cline — an official claim, not our measurement, and not a reply to that thread. Also check the generation on any price you read: a $10/month review online refers to the GLM-5.1 era (March 2026), other pages list "$18/$72/$160" caps or call $18 an "alternative" price. Our numbers above come from the official subscribe page on : Lite is $18/mo, $12.60/mo billed yearly. Dates and generations move faster than the articles about them — including this one.
Checked . Plan prices from z.ai/subscribe (monthly and yearly billing); credit allowances, multipliers, and routing rules from docs.z.ai devpack documentation; campaign details from the official GLM-5.3-Flash Usage Campaign notice. All times UTC+8 per the official notice.
Sources: Z.ai GLM Coding Plan subscription page · Z.ai devpack documentation · GLM-5.3-Flash Usage Campaign notice · r/ZaiGLM thread (community quality debate)