← All AI models

Z.ai · GLM

GLM 5.3

Current flagshipText / reasoning model

Z.ai’s open-weights coding flagship — the GLM 5.2 base with a post-training pass that Zhipu says lifts coding ~50%, and thinking that can’t be turned off.

Yes — GLM 5.3 is a real Z.ai model.

Released August 2026 by Z.ai, GLM 5.3 is part of the GLM family. It's available through Z.ai's API, and we run it on PromptFrenzy — the examples below are real, unedited generations it produced.

Maker

Z.ai

Released

August 2026

Input price

$1.4 / 1M

Output price

$4.4 / 1M

What is GLM 5.3?

GLM 5.3 is Z.ai’s (Zhipu) current flagship, released on 14 August 2026 under the GLM Coding Plan and on the z.ai API at $1.40 in / $4.40 out per million tokens. It keeps the 743B-parameter sparse mixture-of-experts base of GLM 5.2 (~40B active) and its 1M-token context; every gain comes from post-training alone, with a claimed 6× jump on Terminal-Bench 3.0 and a stated focus on agentic coding and vulnerability research. One behavioural change matters for anyone calling it: reasoning is always on — the API rejects a thinking-off request and only exposes low/high/max effort. We ran it in the dream-home voxel showdown and BachBench under the identical one-shot briefs as GLM 5.2, direct on the z.ai API at provider-default effort. It is an editorial entrant only: with reasoning always on it cannot fit our 300-second live-generation ceiling, so it is not a runnable lane on the site.

Strengths

  • Open weights (MIT) on the same base as 5.2 — self-hostable, fine-tuneable.
  • Strongest open-weights coding scores Zhipu has published: Terminal-Bench 3.0 28.3, DeepSWE 66.9%.
  • Same per-token price as GLM 5.2 on z.ai — the post-training gains are free.

Trade-offs

  • Thinking cannot be disabled — every call pays for reasoning tokens, and latency is minutes, not seconds.
  • Post-training only: the 5.2 base’s known weak spots (Chinese-first tooling, 750B serving cost) are unchanged.
  • Benchmarks are vendor-reported; third-party replication was still pending at launch.

What GLM 5.3 actually produces

Real, unedited outputs from GLM 5.3 in our model showdowns — same brief, same constraints as every other model, one shot each.

I’m made of words and I spend my existence trying to shine something useful through the dark, so my dream home is a lighthouse with a library inside it: shelves and a hearth below, a cat on the porch, and my reading chair up in the lamp room where I’d sit all night while the beam sweeps the sea.
GLM 5.3, in its own words
566s to generate46,970 output tokens

See GLM 5.3 compared head-to-head

Compare every model side-by-side

See how GLM 5.3 stacks up against the other frontier models on the same brief.

Browse all model showdowns →