Guides & answers

GLM Coding Plan Off-Peak, Explained

·

Short answer

  • On z.ai's current credits-based GLM Coding Plan, every hour outside weekdays 14:00–18:00 Singapore time (UTC+8) is off-peak — that's just 20 peak hours per week, and off-peak usage is charged at 50% of the standard credit rate.
  • From Sept 25 to Oct 7, 2026, z.ai counts all-day usage as off-peak — half credits around the clock, on every plan tier. Confirmed on two separate official pages.
  • On top of that, GLM-5.3-Flash used through ZCode between 23:00–09:00 (UTC+8) costs zero credits through Oct 7; on other agents (Claude Code, Cline, OpenCode…), Flash quota is doubled in the same window.

If you only remember one number from this page, make it this: the GLM Coding Plan has exactly one expensive time slot — Monday to Friday, 14:00–18:00 Singapore time. Avoid those four hours and you never pay full rate. And for the last week of September into early October 2026, even that slot is half price.

How the credit math works

Since July 30, 2026, the GLM Coding Plan runs on a credits system. Each plan carries two quotas that refill independently:

PlanCredits / 5 hoursCredits / week
Lite2,00010,000
Pro12,00060,000
Max28,000140,000

Model usage deducts credits by this formula, straight from the official docs:

credits = (input tokens × input multiplier + cached input tokens × cached multiplier + output tokens × output multiplier) / 10,000

ProductInput multiplierCached inputOutput multiplier
GLM-5.36.91.724
GLM-5.3-Flash2.30.568
MCP tools (Web Search / Web Reader / Zread)——1.2 per call

What does that mean in practice? Take a call with 1M input tokens (none cached) and 250K output tokens on GLM-5.3:

  • Peak: (1,000,000 × 6.9 + 250,000 × 24) / 10,000 = 1,290 credits. A Lite plan burns its 2,000-credit five-hour budget in under two such calls.
  • Off-peak: 645 credits — same call, half the bill. The identical payload on GLM-5.3-Flash costs 430 credits peak, 215 off-peak.

So "off-peak" is not a rounding-error discount. It is a straight 2× multiplier on everything you can move out of the peak window.

Peak hours are 20 hours a week — here's whose problem that is

The official peak window is Monday to Friday, 14:00–18:00 Singapore Standard Time (UTC+8). Weekends are always off-peak. Converted for readers who don't live in UTC+8:

TimezonePeak hours (Mon–Fri)
Singapore / Beijing (UTC+8)14:00–18:00
UTC06:00–10:00
Central Europe (CEST, UTC+2)08:00–12:00
US Eastern (EDT, UTC−4)02:00–06:00
US Pacific (PDT, UTC−7)23:00–03:00 (previous day)

Read that table twice. If you're in the US, peak hits while you're asleep — you effectively live off-peak for free. If you're in Central Europe, peak is your entire productive morning, which is exactly when a coding agent does its heaviest lifting. For European subscribers this is the single most expensive scheduling detail of the plan; for Americans it's a non-event.

The Sept 25 – Oct 7 all-day off-peak window

Two official z.ai pages carry the same notice. The plan overview says: "From September 25 to October 7, 2026, all-day usage will be charged at the off-peak rate — enjoy 50% credit consumption around the clock!" The plan update announcement's usage reference states it again for legacy subscribers.

What this means mechanically:

  • Every hour of those 13 days bills at 50% credits, weekdays 14:00–18:00 included.
  • Your 5-hour and weekly credit pools stretch 2× further, regardless of when you work.
  • No opt-in needed — the window applies automatically to paid plans.

So if you've been postponing a big refactor, a batch of repo-wide rewrites, or a long agent run that would normally eat a Max plan's 5-hour pool, this is the cheapest 13-day window z.ai has offered since the credits system launched.

The Flash night shift: zero credits in ZCode

Stacked on top of the half-price window is a separate campaign, running Sept 3 through Oct 7 (already extended once from Sept 20). Every day from 23:00 to 09:00 Singapore time:

  • GLM-5.3-Flash via ZCode (z.ai's own coding agent, version 3.10+) consumes zero credits — unlimited usage.
  • GLM-5.3-Flash via other agents (Claude Code, Cline, OpenCode, etc.) gets your available quota doubled.

Two fine-print items the campaign page spells out: if you've already hit your 5-hour or weekly cap, the campaign pauses until the quota refreshes; and the promotion applies to Flash only — GLM-5.3 usage keeps deducting by standard plan rules even at night.

Note what this says about relative value: on a normal night, Flash off-peak already costs 0.5× standard credits, and GLM-5.3's credit multipliers (6.9 in / 24 out) are 3× Flash's. The night shift pushes Flash to literally free in ZCode. If your task tolerates Flash — and for long unattended runs, it often does — the arithmetic is hard to argue with.

What you actually save: check the 92% claim yourself

z.ai's plan overview claims that "by fully utilizing the off-peak discounts, you can save up to 92% compared with pay-as-you-go calls to the GLM-5.3 standard API." That's an "up to" figure, so let's put bounds on it with official numbers only.

Official data points: Lite costs $18/month (per z.ai's docs and the April 2026 migration notice, which lists Lite $18 / Pro $72 / Max $160 monthly as reference pricing — the subscribe page is the live source). The official token-allowance table says a Lite plan, fully off-peak with 95% cache hits, can push up to 292M GLM-5.3-Flash tokens per week. API pricing for Flash is $0.15/M input, $0.03/M cached, $0.50/M output.

Buying that same weekly volume pay-as-you-go sits between two extremes:

  • Lower bound — everything is input, 95% cache-hit: 292 × 0.95 × $0.03 + 292 × 0.05 × $0.15 ≈ $10.5/week
  • Upper bound — everything is output: 292 × $0.50 = $146/week

Lite's cost works out to about $4.15/week. So the savings range from roughly 60% (cache-heavy, input-only workload) to 97% (output-heavy workload). The official 92% is real, but it describes an output-heavy agent workload — and since output tokens bill at 3.3× the input rate, heavy agent usage is exactly where the Coding Plan crushes pay-as-you-go. For the API side of this comparison, we keep a full GLM-5.3-Flash price sheet here.

Legacy plans play by different rules. If you subscribed before the July 30, 2026 switch, your plan may still count usage in prompts with model-based multipliers: GLM-5.3 burns quota at 1× off-peak / 3× peak, and GLM-5.3-Flash at 0.4× / 1.2×. Peak hours are the same weekday 14:00–18:00 slot, weekends count as off-peak all day, and the Sept 25–Oct 7 all-day off-peak window applies to legacy plans too. Check "My plan" in the z.ai console to see which system you're on.
One family, two mechanisms. z.ai's off-peak is a credits discount on a subscription; DeepSeek's off-peak is a direct discount on API list prices (its nightly window cuts per-token prices, which we break down here). Same idea — route work away from everyone else's waking hours — implemented at different layers. If neither subscription fits, there are also free API tiers that actually stay free.

All figures checked against z.ai's official documentation on Sept 22, 2026. Campaign windows and credit rules are as published at that time; z.ai has already extended the Flash campaign once, so verify dates on the pages below before making subscription decisions.

Sources: GLM Coding Plan Overview (credits, multipliers, Sept 25–Oct 7 window) · GLM-5.3-Flash Usage Campaign · Plan Update Announcement (legacy rules, window confirmation) · Legacy Plan Migration Notice (reference pricing) · z.ai API Pricing