# GLM-4.6V-Flash API prices

GLM-4.6V-Flash (glm-flash) — 200k context, 128k max output.
1 metered per-token offers from 6 provider(s), in USD per 1M tokens.
Prices as of 2026-10-04.

| Provider | Input $/1M | Output $/1M | Cache read $/1M | Context |
| --- | --- | --- | --- | --- |
| Hugging Face | $0.3 | $0.9 | — | 131k |

## Summary

- Cheapest input: $0.3 per 1M tokens (Hugging Face)
- Cheapest output: $0.9 per 1M tokens (Hugging Face)
- First-party: free per 1M output tokens (Z.AI)
- Free offers: EmpirioLabs AI, Tempr Gateway, Z.AI, ZenMux, Zhipu AI
- Subscription plans (not per-token): none
- Released: 2025-12-08
- Inputs: image, pdf, text, video

## FAQ

### What is the cheapest GLM-4.6V-Flash API?

As of Oct 4, 2026, Hugging Face has the lowest GLM-4.6V-Flash output price at $0.9 per 1M tokens, and Hugging Face has the lowest input price at $0.3 per 1M tokens. Prices are metered per-token rates in USD; free tiers and subscription plans are excluded.

### How much does GLM-4.6V-Flash cost on Z.AI?

Z.AI charges free per 1M input tokens and free per 1M output tokens for GLM-4.6V-Flash.

### How many providers offer GLM-4.6V-Flash?

6 providers list GLM-4.6V-Flash on Sovyron; 1 of them sell it at a metered per-token price.

### Is GLM-4.6V-Flash free?

5 provider(s) list a free-tier offer: EmpirioLabs AI, Tempr Gateway, Z.AI, ZenMux, Zhipu AI. Free tiers usually have rate limits.

### What is the context window of GLM-4.6V-Flash?

GLM-4.6V-Flash supports a 200k-token context window and up to 128k output tokens.

HTML page: https://sovyron.com/models/glm-4-6v-flash/
Full dataset: https://sovyron.com/data/catalog.json
