Skip to content
Sovyron

GLM-4.6V-Flash

glm-flash200k context128k max outputreleased Dec 8, 2025Open weightsReasoningTool calling

As of Oct 4, 2026, the cheapest GLM-4.6V-Flash API price is $0.3 per 1M input tokens (Hugging Face) and $0.9 per 1M output tokens (Hugging Face). GLM-4.6V-Flash is offered by 6 providers in the glm-flash family, open weights, reasoning enabled. Z.AI's own rate is free per 1M output tokens. 5 providers list it as a free-tier offer.

Cheapest input
$0.3
Cheapest output
$0.9
Official (Z.AI)
free
output per 1M tokens
Providers
6
offering this model
Spread
—
max / min output
Free offers
5
free-tier offers

API pricing by provider

Metered per-token prices. Cheapest input and output are highlighted; † marks a context-tier surcharge. How prices are ranked.

GLM-4.6V-Flash API prices by provider in USD per 1M tokens, updated Oct 4, 2026
Hugging FaceopenCheapest$0.3Cheapest$0.9——131k
EmpirioLabs AIfreefreefree——128k
Tempr Gatewayfreefreefreefreefree128k
Z.AIfreefreefreefreefree128k
ZenMuxfreefreefreefree—200k
Zhipu AIfreefreefreefreefree128k

GLM-4.6V-Flash pricing FAQ

What is the cheapest GLM-4.6V-Flash API?

As of Oct 4, 2026, Hugging Face has the lowest GLM-4.6V-Flash output price at $0.9 per 1M tokens, and Hugging Face has the lowest input price at $0.3 per 1M tokens. Prices are metered per-token rates in USD; free tiers and subscription plans are excluded.

How much does GLM-4.6V-Flash cost on Z.AI?

Z.AI charges free per 1M input tokens and free per 1M output tokens for GLM-4.6V-Flash.

How many providers offer GLM-4.6V-Flash?

6 providers list GLM-4.6V-Flash on Sovyron; 1 of them sell it at a metered per-token price.

Is GLM-4.6V-Flash free?

5 provider(s) list a free-tier offer: EmpirioLabs AI, Tempr Gateway, Z.AI, ZenMux, Zhipu AI. Free tiers usually have rate limits.

What is the context window of GLM-4.6V-Flash?

GLM-4.6V-Flash supports a 200k-token context window and up to 128k output tokens.

Token cost calculator

What a month costs at one provider's published rates. Enter the tokens you expect to send and receive; cached tokens use the cached-input price where the provider lists one, and count as ordinary input where it does not.

Estimated

—