Skip to content
Sovyron

GLM-5.3-Flash

glm-flash1.0M context1.0M max outputreleased Aug 26, 2026Open weightsReasoningTool calling4 subscription plans →

As of Oct 4, 2026, the cheapest GLM-5.3-Flash API price is $0.03 per 1M input tokens (Pareto Inference) and $0.025 per 1M output tokens (TokenGo). GLM-5.3-Flash is offered by 69 providers in the glm-flash family, open weights, reasoning enabled. Z.AI's own rate is $0.5 per 1M output tokens. 4 providers list it as a free-tier offer. It is also included in 4 flat-rate subscription plans, OpenCode OpenCode Go from $10/mo.

Cheapest input
$0.03
Cheapest output
$0.025
Official (Z.AI)
$0.5
output per 1M tokens
Providers
69
offering this model
Spread
60.0×
max / min output
Free offers
4
free-tier offers

API pricing by provider

Metered per-token prices. Cheapest input and output are highlighted; † marks a context-tier surcharge. How prices are ranked.

GLM-5.3-Flash API prices by provider in USD per 1M tokens, updated Oct 4, 2026
Pareto InferenceopenCheapest$0.03$0.1$0.006—1.0M—
TokenGoopen$0.075Cheapest$0.025$0.015—1.0M—
CrofAIopen$0.07$0.22$0.01—1.0M—
302.AIopen$0.075$0.25——1.0M—
EmpirioLabs AIopen$0.075$0.25$0.075—1.0M—
Merge Gatewayopen$0.075$0.25$0.015—1.0M3d agoout +400%
OrcaRouteropen$0.075$0.25——1.0M—
DevPass (LLM Gateway)open$0.088$0.25$0.025—1.0M10d agoout +31.6%
LLM Gatewayopen$0.088$0.25$0.025—1.0M—
LLM Gatewayopen$0.09$0.28$0.02—1.0M—
NanoGPTopen$0.1$0.3$0.025—1.0M3d agoout +20%
Cortecsopen$0.1$0.35$0.018—1.0M22d agoout −30%
Vultropen$0.1$0.35——1.0M—
RunInfraopen$0.1$0.4$0.01—1.0M—
AIHubMixopen$0.1127$0.3944$0.0282—1.0M—
Vancineopen$0.12$0.4$0.024—1.0M19d agoout +100%
Volcengine Arkopen$0.1187$0.4156$0.0341—1.0M—
Bothubopen$0.12$0.44——1.0M—
Meliousopen$0.1159$0.4637$0.0232—1.0M—
engyopen$0.135$0.45$0.027—262k—
GreenPTopen$0.1278$0.511$0.0256—1.0M—
IteraComputeopen$0.14$0.49$0.03—1.0M—
ai&open$0.15$0.5$0.03—1.0M—
Basetenopen$0.15$0.5——1.0M—
Cloudflare Workers AIopen$0.15$0.5$0.03—1.0M—
CrossModelopen$0.15$0.5$0.03$0.151.0M—
Deep Infraopen$0.15$0.5$0.03—1.0M—
Eden AIopen$0.15$0.5$0.03—1.0M—
Fireworks AIopen$0.15$0.5$0.03—1.0M—
Friendliopen$0.15$0.5$0.03—1.0M—
Hugging Faceopen$0.15$0.5——1.0M—
Incoopen$0.15$0.5——1.0M—
Kilo Gatewayopen$0.15$0.5$0.03—1.0M—
LLM Gatewayopen$0.15$0.5$0.03—1.0M—
LLM Gatewayopen$0.15$0.5$0.03—1.0M—
LLM Gatewayopen$0.15$0.5$0.03—1.0M—
LLM Gatewayopen$0.15$0.5$0.03—1.0M—
Nebius Token Factoryopen$0.15$0.5$0.15—1.0M—
Neonopen$0.15$0.5$0.03—1.0M—
Neuralwattopen$0.15$0.5$0.03—1.0M—
Ofoxopen$0.15$0.5$0.03—1.0M—
Ollama Cloudopen$0.15$0.5$0.03—1.0M—
OpenCode Zenopen$0.15$0.5$0.03—1.0M—
OpenCode Goopen$0.15$0.5$0.03—1.0M—
OpenRouteropen$0.15$0.5$0.03—1.0M6d agoout +257.1%
SiliconFlowopen$0.15$0.5$0.03free1.0M—
Syntheticopen$0.15$0.5$0.04—524k—
Tempr Gatewayopen$0.15$0.5$0.03free1.0M—
Together AIopen$0.15$0.5$0.03—1.0M—
Umans AIopen$0.15$0.5$0.03—1.0M—
Venice AIopen$0.15$0.5$0.03—1.0M—
Vercel AI Gatewayopen$0.15$0.5$0.03—1.0M—
Vivgridopen$0.15$0.5$0.04—1.0M—
CoreWeaveopen$0.15$0.5$0.05—1.0M—
Z.AIopen$0.15$0.5$0.03free1.0M15d agoout +100%
ZenMuxopen$0.15$0.5$0.03—1.0M—
Zhipu AIopen$0.15$0.5$0.03free1.0M15d agoout +100%
Charm Hyperopen$0.1633$0.5444$0.0316—1.0M—
above.devopen$0.165$0.55$0.0319—1.0M—
Opperopen$0.2$0.5$0.07—1.0M—
TensorXopen$0.2$0.5$0.05—1.0M—
Requestyopen$0.2$0.6$0.07—1.0M—
Requestyeuopen$0.2$0.6$0.07—1.0M—
Privatemode AIopen$0.2311$0.7511$0.0578—1.0M6d agoout −83.2%
Berget.AIopen$0.29$0.58——524k—
Tinfoilopen$0.4$1.25$0.1—1.0M—
Modalopen$0.45$1.5$0.09—1.0M—
Kenarifreefreefree——1.0M—
NaNfreefreefree——1.0M—
Nvidiafreefreefree——1.0M—
OrcaRouterfreefreefree——1.0M—

Subscription plans (not per-token)

Billed per month, not per token — never counted as the cheapest offer.

SCNet Token Planplanfreefree1.0M
Umans AI Coding Planplanfreefree1.0M
Volcengine Ark Coding Planplanfreefree1.0M
Z.AI Coding Planplanfreefree1.0M
Zhipu AI Coding Planplanfreefree1.0M

GLM-5.3-Flash pricing FAQ

What is the cheapest GLM-5.3-Flash API?

As of Oct 4, 2026, TokenGo has the lowest GLM-5.3-Flash output price at $0.025 per 1M tokens, and Pareto Inference has the lowest input price at $0.03 per 1M tokens. Prices are metered per-token rates in USD; free tiers and subscription plans are excluded.

How much does GLM-5.3-Flash cost on Z.AI?

Z.AI charges $0.15 per 1M input tokens and $0.5 per 1M output tokens for GLM-5.3-Flash.

How many providers offer GLM-5.3-Flash?

69 providers list GLM-5.3-Flash on Sovyron; 67 of them sell it at a metered per-token price.

Is GLM-5.3-Flash free?

4 provider(s) list a free-tier offer: Kenari, NaN, Nvidia, OrcaRouter. Free tiers usually have rate limits.

Is GLM-5.3-Flash included in a subscription plan?

Yes. 4 flat-rate plan(s) include it; the cheapest is OpenCode OpenCode Go at $10/month.

What is the context window of GLM-5.3-Flash?

GLM-5.3-Flash supports a 1.0M-token context window and up to 1.0M output tokens.

Token cost calculator

What a month costs at one provider's published rates. Enter the tokens you expect to send and receive; cached tokens use the cached-input price where the provider lists one, and count as ordinary input where it does not.

Estimated

—