Nebius Token Factory
21 modelsProvider docs
As of Oct 4, 2026, Nebius Token Factory lists 21 supported models in the Sovyron catalog, including DeepSeek V4.1 Flash, GLM-5.3-Flash, GLM-5.3, Qwen3.8 27B and DeepSeek V4 Pro 0813 — deepseek-flash, glm-flash, glm, qwen, deepseek-thinking and nemotron families. Its metered rates span $0.0075–$6 per 1M tokens on a 3:1 input:output blend, and every price in the table is Nebius Token Factory's own, normalized to USD — open a model to compare it with the cheapest offer across all providers.
Supported models and pricing
21 of 21 models
| Model | 1M input | 1M output | 1M cache read | Context | Released |
|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | $0.3 | $1.2 | $0.3 | 1.0M | Sep 10, 2026 |
| GLM-5.3-Flash | $0.15 | $0.5 | $0.15 | 1.0M | Aug 26, 2026 |
| GLM-5.3 | $1.4 | $4.4 | $1.4 | 1.0M | Aug 14, 2026 |
| Qwen3.8 27B | $0.45 | $3 | $0.45 | 262k | Aug 14, 2026 |
| DeepSeek V4 Pro 0813 | $1.32 | $3.96 | $1.32 | 979k | Aug 12, 2026 |
| Nemotron 3.5 Lightning 30B A3B | $0.06 | $0.24 | $0.06 | 1.0M | Aug 11, 2026 |
| DeepSeek V4 Flash 0731 | $0.14 | $0.28 | $0.14 | 1.0M | Jul 31, 2026 |
| Kimi K3 | $3 | $15 | $3 | 1.0M | Jul 16, 2026 |
| GLM-5.2 | $1.4 | $4.4 | — | 1.0M | Jun 13, 2026 |
| Kimi K2.7 Code | $0.95 | $4 | — | 262k | Jun 12, 2026 |
| Nemotron 3 Ultra 550B A55B | $1 | $3 | $1 | 1.0M | Jun 4, 2026 |
| MiniMax-M3 | $0.3 | $1.2 | — | 1.0M | Jun 1, 2026 |
| DeepSeek V4 Pro | $1.75 | $3.5 | $0.15 | 1.0M | Apr 24, 2026 |
| Nemotron 3 Super 120B A12B | $0.3 | $0.9 | — | 262k | Mar 11, 2026 |
| Hermes 4 405B | $1 | $3 | $0.1 | 131k | Jan 30, 2026 |
| Qwen3 30B A3B Instruct 2507 | $0.1 | $0.3 | $0.01 | 262k | Jan 28, 2026 |
| Gemma 3 27B IT | $0.1 | $0.3 | $0.01 | 110k | Jan 20, 2026 |
| GPT OSS 120B | $0.15 | $0.6 | $0.015 | 131k | Jan 10, 2026 |
| Qwen3 Embedding 8B | $0.01 | free | — | 41k | Jan 10, 2026 |
| Qwen3 235B A22B Instruct 2507 | $0.2 | $0.6 | — | 262k | Jul 25, 2025 |
| Qwen3.5 397B-A17B | $0.6 | $3.6 | $0.06 | 262k | Jul 15, 2025 |