Llama 3.2 3B Instruct
As of Oct 4, 2026, the cheapest Llama 3.2 3B Instruct API price is $0.02 per 1M input tokens (Inference) and $0.02 per 1M output tokens (Inference). Llama 3.2 3B Instruct is offered by 9 providers in the llama family, open weights, reasoning enabled.
API pricing by provider
Metered per-token prices. Cheapest input and output are highlighted; † marks a context-tier surcharge. How prices are ranked.
| Inferenceopen | Cheapest$0.02 | Cheapest$0.02 | — | — | 16k |
| DevPass (LLM Gateway)open | $0.03 | $0.05 | — | — | 33k |
| LLM Gateway | $0.03 | $0.05 | — | — | 33k |
| NovitaAIopen | $0.03 | $0.05 | — | — | 33k |
| NanoGPTopen | $0.0306 | $0.0493 | $0.0153 | — | 131k |
| Kilo Gateway | $0.05 | $0.33 | — | — | 131k |
| OpenRouteropen | $0.05 | $0.33 | — | — | 131k |
| Cloudflare Workers AIopen | $0.0509 | $0.335 | — | — | 80k |
| Pioneeropen | $0.1 | $0.335 | $0.1 | $0.1 | 131k |
Llama 3.2 3B Instruct pricing FAQ
What is the cheapest Llama 3.2 3B Instruct API?
As of Oct 4, 2026, Inference has the lowest Llama 3.2 3B Instruct output price at $0.02 per 1M tokens, and Inference has the lowest input price at $0.02 per 1M tokens. Prices are metered per-token rates in USD; free tiers and subscription plans are excluded.
How many providers offer Llama 3.2 3B Instruct?
9 providers list Llama 3.2 3B Instruct on Sovyron; 9 of them sell it at a metered per-token price.
Is Llama 3.2 3B Instruct free?
No provider in the Sovyron catalog lists a free tier for Llama 3.2 3B Instruct.
What is the context window of Llama 3.2 3B Instruct?
Llama 3.2 3B Instruct supports a 131k-token context window and up to 118k output tokens.
Token cost calculator
What a month costs at one provider's published rates. Enter the tokens you expect to send and receive; cached tokens use the cached-input price where the provider lists one, and count as ordinary input where it does not.
Estimated
—