Z-AI

GLM 5.3 Flash

z-ai/glm-5.3-flashHugging Face ↗

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

toolsstructuredreasoningimage in
Data source

Only tracked on Vercel AI Gateway.

GLM 5.3 Flash usage over time

Vercel AI Gateway share of tokens, by day.

View
Tidelines — LLM usage and market share

GLM 5.3 Flash holds 6.91% today, up from 1.78% at the start of the visible window.

Data: Vercel AI Gateway as of 2026-09-02 · Source: tidelines.ai/models/z-ai-glm-5-3-flash

Rank history

Tidelines — LLM usage and market share
Data: Vercel AI Gateway as of 2026-09-02 · Source: tidelines.ai/models/z-ai-glm-5-3-flash

Adoption curve vs the pantheon

GLM 5.3 Flash's share plotted against days since first tracked, overlaid on other current leaders on the same axis. The question every new-model page asks: is this one going to make it?

Tidelines — LLM usage and market share
Data: Vercel AI Gateway as of 2026-09-02 · Source: tidelines.ai/models/z-ai-glm-5-3-flash

Citable stat

As of 2026-09-02, GLM 5.3 Flash holds 6.9% of Vercel AI Gateway token share, up 6.91pp over 30 days. Source: Tidelines (Vercel AI Gateway).

Cost per session

Median dollars spent on one GLM 5.3 Flash session, by coding agent and session length.

Tidelines — LLM usage and market share

Token rate sets the price; session length and reasoning set the bill.

Agent1 turn2–9 turns10–49 turns50+ turns
Claude Code$0.0022$0.0052$0.038$0.343
Kilo Code$0.0011$0.0059$0.030$0.255
Hermes Agent$0.0015$0.0040$0.028$0.205
Codex$0.0009$0.037

Compare every model's session cost →

Data: OpenRouter session-cost dataset as of 2026-08-30 · Source: tidelines.ai/models/z-ai-glm-5-3-flash

Where to run it

GLM 5.3 Flash pricing by provider

Tidelines — LLM usage and market share

Per-provider pricing, context, and uptime via OpenRouter. Sorted by output price.

GLM 5.3 Flash is served by 23 tracked endpoints on OpenRouter. The cheapest output price is $0.25 per million tokens from GMICloud, and the most expensive is $0.50 from Cloudflare. The largest context window offered is 1311k tokens. Prices are list rates and refresh daily.

ProviderInput $/MOutput $/MContextUptime
GMICloud$0.07$0.251049k96.80%
Novita$0.07$0.251049k99.08%
DeepInfra$0.07$0.251049k99.87%
Z.AI$0.07$0.251049k97.86%
Wafer$0.10$0.351049k95.34%
Morph$0.13$0.451049k99.66%
Makora$0.14$0.471049k98.93%
Modal$0.15$0.501049k99.97%
Sail Research$0.15$0.501049k99.76%
StreamLake$0.15$0.501024k98.75%
NextBit$0.15$0.501049k97.44%
Fireworks$0.15$0.501049k99.53%
Phala$0.15$0.501049k98.15%
Friendli$0.15$0.501049k99.86%
SiliconFlow$0.15$0.501049k97.71%
DigitalOcean$0.15$0.501049k98.40%
Together$0.15$0.501049k99.55%
Reka$0.15$0.50262k99.96%
Parasail$0.15$0.501049k98.30%
BaseTen$0.15$0.501049k99.27%
Venice$0.15$0.501049k99.37%
Io Net$0.15$0.50262k91.34%
Cloudflare$0.15$0.501311k99.92%