Llama 4 Scout 17B 16E Instruct
As of Oct 4, 2026, the cheapest Llama 4 Scout 17B 16E Instruct API price is $0.2 per 1M input tokens (Azure) and $0.78 per 1M output tokens (Azure). Llama 4 Scout 17B 16E Instruct is offered by 3 providers in the llama family, open weights.
API pricing by provider
Metered per-token prices. Cheapest input and output are highlighted; † marks a context-tier surcharge. How prices are ranked.
| Provider | 1M input | 1M output | 1M cache read | 1M cache write | Context |
|---|---|---|---|---|---|
| Azureopen | Cheapest$0.2 | Cheapest$0.78 | — | — | 128k |
| Azure Cognitive Servicesopen | $0.2 | $0.78 | — | — | 128k |
| Cloudflare Workers AIopen | $0.27 | $0.85 | — | — | 131k |
Llama 4 Scout 17B 16E Instruct pricing FAQ
What is the cheapest Llama 4 Scout 17B 16E Instruct API?
As of Oct 4, 2026, Azure has the lowest Llama 4 Scout 17B 16E Instruct output price at $0.78 per 1M tokens, and Azure has the lowest input price at $0.2 per 1M tokens. Prices are metered per-token rates in USD; free tiers and subscription plans are excluded.
How many providers offer Llama 4 Scout 17B 16E Instruct?
3 providers list Llama 4 Scout 17B 16E Instruct on Sovyron; 3 of them sell it at a metered per-token price.
Is Llama 4 Scout 17B 16E Instruct free?
No provider in the Sovyron catalog lists a free tier for Llama 4 Scout 17B 16E Instruct.
What is the context window of Llama 4 Scout 17B 16E Instruct?
Llama 4 Scout 17B 16E Instruct supports a 131k-token context window and up to 16k output tokens.