Skip to content
Sovyron

Llama 4 Maverick 17B 128E Instruct FP8

llama1.0M context16k max outputreleased Apr 5, 2025Open weightsTool calling

As of Oct 4, 2026, the cheapest Llama 4 Maverick 17B 128E Instruct FP8 API price is $0.25 per 1M input tokens (Azure) and $1 per 1M output tokens (Azure). Llama 4 Maverick 17B 128E Instruct FP8 is offered by 5 providers in the llama family, open weights. 2 providers list it as a free-tier offer.

Cheapest input
$0.25
Cheapest output
$1
Official
—
output per 1M tokens
Providers
5
offering this model
Spread
1.5×
max / min output
Free offers
2
free-tier offers

API pricing by provider

Metered per-token prices. Cheapest input and output are highlighted; † marks a context-tier surcharge. How prices are ranked.

Llama 4 Maverick 17B 128E Instruct FP8 API prices by provider in USD per 1M tokens, updated Oct 4, 2026
AzureopenCheapest$0.25Cheapest$1——1.0M
Azure Cognitive Servicesopen$0.25$1——1.0M
watsonx.aiopen$0.371$1.484——131k
Llamafreefreefree——128k
Vercel AI Gatewayfreefreefree——128k

Llama 4 Maverick 17B 128E Instruct FP8 pricing FAQ

What is the cheapest Llama 4 Maverick 17B 128E Instruct FP8 API?

As of Oct 4, 2026, Azure has the lowest Llama 4 Maverick 17B 128E Instruct FP8 output price at $1 per 1M tokens, and Azure has the lowest input price at $0.25 per 1M tokens. Prices are metered per-token rates in USD; free tiers and subscription plans are excluded.

How many providers offer Llama 4 Maverick 17B 128E Instruct FP8?

5 providers list Llama 4 Maverick 17B 128E Instruct FP8 on Sovyron; 3 of them sell it at a metered per-token price.

Is Llama 4 Maverick 17B 128E Instruct FP8 free?

2 provider(s) list a free-tier offer: Llama, Vercel AI Gateway. Free tiers usually have rate limits.

What is the context window of Llama 4 Maverick 17B 128E Instruct FP8?

Llama 4 Maverick 17B 128E Instruct FP8 supports a 1.0M-token context window and up to 16k output tokens.

Token cost calculator

What a month costs at one provider's published rates. Enter the tokens you expect to send and receive; cached tokens use the cached-input price where the provider lists one, and count as ordinary input where it does not.

Estimated

—