AI coding tools

GLM-5.3 on Amazon Bedrock

·

Z.ai's flagship GLM-5.3 — the 753B-parameter model that postured its way to the top of Z.ai's own CyberGym security benchmark — is now generally available on Amazon Bedrock. The AWS announcement landed in early October 2026, and most coverage stopped at "model now available." That skips the part an actual buyer cares about: Bedrock's on-demand price for GLM-5.3 is a flat 20% above Z.ai's direct API price, and the Flex tier undercuts direct pricing by 40% if your workload tolerates slower queueing. Here is the full math, plus the access restrictions AWS buries in the model card.

Short answer

  • GLM-5.3 on Bedrock costs $1.68/$5.28 per million input/output tokens through the Global cross-Region inference profile — exactly 1.2× Z.ai's direct API price of $1.40/$4.40.
  • The Flex service tier flips the math: at 50% off Standard, Bedrock charges $0.84/$2.64 — cheaper than direct. Priority tier runs a 75% premium at $2.94/$9.24.
  • Not everyone can buy it: AWS gates GLM-5.3 to "eligible enterprise customers," and the model only runs through cross-Region inference profiles (us.zai.glm-5.3 or global.zai.glm-5.3) — no in-region on-demand calls.
  • The model itself: 753B-parameter MoE with roughly 40B active per token, 1M-token context, 128K max output, reasoning always on with low/high/max effort levels, and prompt caching (explicit cache checkpoints need 1,024 tokens minimum).
  • If you just want the cheapest GLM-5.3 tokens, Z.ai direct still wins for latency-sensitive work — and the smaller GLM-5.3-Flash at $0.15/$0.50 remains the budget workhorse of the family.

What actually shipped

GLM-5.3 is Z.ai's flagship, built on the same base model as GLM-5.2 with improvements driven by scaled post-training. The spec sheet: 753B total parameters in a mixture-of-experts architecture with roughly 40B active per token, a 1M-token context window, and 128K-token maximum output. Text-only input. Reasoning is always enabled — Z.ai removed the ability to turn it off entirely — with three effort levels (low, high, max, defaulting to max). AWS's What's New post describes "selectable effort levels"; the Z.ai docs are more specific, and they include a migration warning worth heeding: applications still sending thinking.type: "disabled" will get request failures until they switch to enabled and dial effort down to low.

On benchmarks, the claims come from Z.ai itself, so weight them accordingly: a 50% coding-performance gain over GLM-5.2 on Z.ai's internal Code Bench, SOTA-among-open-models results on Terminal Bench 3.0 and Agents' Last Exam (CLI), and — the one AWS's own blog chose to highlight — a leading 84.5 on the CyberGym vulnerability-discovery benchmark at release. That security angle is unusual for a general flagship, and AWS's walkthrough demo leans into it, running an authorized penetration test with the open-source Strix agent against a deliberately vulnerable Juice Shop app.

One structural note: this is the second-gen Z.ai model on Bedrock. GLM 5 arrived earlier in 2026, and AWS's blog explicitly says GLM-5.3's benchmark gains were large enough that Z.ai refreshed its benchmark suite rather than compare across generations. Chinese business press (Red Star Capital Bureau, via Sina, October 8) reports AWS pays Zhipu a revenue share based on invocation volume — the standard Bedrock third-party model arrangement.

The Bedrock-specific fine print

This is the section English coverage mostly skipped, and it changes who this launch is actually for:

  • Enterprise gating: "Access to GLM-5.3 on Bedrock is available to eligible enterprise customers" — both AWS's blog and the What's New post repeat this line. It is not a self-serve, any-account model the way Anthropic's or Meta's Bedrock listings are.
  • Cross-Region only: requests must target the US profile (us.zai.glm-5.3) or Global profile (global.zai.glm-5.3). Direct in-region on-demand inference is not supported, which means slightly higher and more variable latency than a single-region deployment.
  • Caching works, and matters: implicit (automatic) prompt caching is on by default, and explicit cache checkpoints are supported with a 1,024-token minimum per checkpoint and at least 30 minutes retention. For agentic loops that resend a large system prompt every turn, this is where the effective cost gap narrows.
  • Service tiers apply: Standard (pay-per-token, no commitment), Priority (fastest, +75%), Flex (cheapest, 50% off, for queue-tolerant workloads), and Reserved throughput for committed volume.
  • API surface: OpenAI-compatible Responses and Chat Completions APIs, plus native Invoke and Converse. AWS's blog demonstrates wiring it into coding assistants like OpenCode.

The price math

Bedrock's published on-demand rates for GLM-5.3, straight from AWS's pricing page (Standard tier):

ChannelInput / 1MOutput / 1MCache read / 1M
Z.ai direct API$1.40$4.40$0.26
Bedrock — Global CRIS$1.68$5.28$0.312
Bedrock — US CRIS$1.848$5.808$0.3432
Bedrock — Global, Flex tier$0.84$2.64—
Bedrock — Global, Priority tier$2.94$9.24—

The pattern is tidy: AWS charges a flat 20% premium over direct pricing on both input and output (1.68/1.40 and 5.28/4.40 both land at 1.2), the US profile adds another 10% over Global, and the Flex tier cuts Standard in half — dropping below Z.ai's own list price. Cache writes on Bedrock Global run $2.10 per million tokens; Z.ai currently lists cached-input storage as free for a limited time.

What that means in practice: a workload burning 100M input + 30M output tokens per month costs $272/month direct versus $326 on Bedrock Standard Global — a $55 premium that buys AWS-billed invoicing, VPC-adjacent architecture, and compliance posture rather than a different model. Run the same volume through Flex and it drops to roughly $163, assuming the workload tolerates flex queueing. For teams already committed to AWS procurement, the premium is frequently cheaper than the internal fight to approve a new vendor.

Which channel, then

Z.ai direct is the rational default if you're a startup or indie team: cheaper on Standard-rate tokens, same model, first-day access to new releases, and the GLM Coding Plan subscription covers GLM-5.3 for coding-assistant use at a flat monthly rate that makes per-token math irrelevant for individuals. Bedrock earns its 20% when enterprise requirements — consolidated AWS billing, security review sign-off, data-residency posturing — are the actual bottleneck. And if your workload is batch-shaped (evaluation runs, bulk summarization, overnight agent sweeps), Bedrock Flex at $0.84/$2.64 is currently the cheapest metered way to run GLM-5.3 anywhere, including direct.

There's a third door worth remembering: the model is open-weight. If your volume is large and steady, self-hosting beats both — our GLM-5.3-Flash self-hosting hardware breakdown walks the economics for the smaller variant, and the same logic scales up with steeper hardware requirements at 753B total parameters.

Sources and caveats

Price snapshot taken October 9, 2026. All Bedrock figures come from AWS's Bedrock pricing page (Standard tier, On-Demand, GLM 5.3 section); Z.ai figures from Z.ai's official pricing documentation. Model specifications and benchmark claims are from AWS's What's New announcement, the AWS Machine Learning blog, and Z.ai's developer documentation — benchmark numbers (CyberGym 84.5, the 50% Code Bench gain) are Z.ai's own reports, not independently verified. The "eligible enterprise customers" language is quoted from AWS's posts; the exact eligibility criteria are not public. Preview-era pricing changes without much ceremony — re-check both price pages before committing a budget. No affiliate links; Z.ai and AWS did not review or pay for this page.