Skip to content
Sovyron

Llama 3.2 1B Instruct

llama131k context60k max outputreleased Sep 25, 2024Open weightsReasoningTool calling

As of Oct 4, 2026, the cheapest Llama 3.2 1B Instruct API price is $0.01 per 1M input tokens (Inference) and $0.01 per 1M output tokens (Inference). Llama 3.2 1B Instruct is offered by 5 providers in the llama family, open weights, reasoning enabled.

Cheapest input
$0.01
Cheapest output
$0.01
Official
—
output per 1M tokens
Providers
5
offering this model
Spread
20.1×
max / min output
Free offers
0
free-tier offers

API pricing by provider

Metered per-token prices. Cheapest input and output are highlighted; † marks a context-tier surcharge. How prices are ranked.

Llama 3.2 1B Instruct API prices by provider in USD per 1M tokens, updated Oct 4, 2026
InferenceopenCheapest$0.01Cheapest$0.01——16k
Cloudflare Workers AIopen$0.027$0.201——60k
Kilo Gateway$0.027$0.201——60k
OpenRouteropen$0.027$0.201——60k
Pioneeropen$0.1$0.201$0.1$0.1131k

Llama 3.2 1B Instruct pricing FAQ

What is the cheapest Llama 3.2 1B Instruct API?

As of Oct 4, 2026, Inference has the lowest Llama 3.2 1B Instruct output price at $0.01 per 1M tokens, and Inference has the lowest input price at $0.01 per 1M tokens. Prices are metered per-token rates in USD; free tiers and subscription plans are excluded.

How many providers offer Llama 3.2 1B Instruct?

5 providers list Llama 3.2 1B Instruct on Sovyron; 5 of them sell it at a metered per-token price.

Is Llama 3.2 1B Instruct free?

No provider in the Sovyron catalog lists a free tier for Llama 3.2 1B Instruct.

What is the context window of Llama 3.2 1B Instruct?

Llama 3.2 1B Instruct supports a 131k-token context window and up to 60k output tokens.

Token cost calculator

What a month costs at one provider's published rates. Enter the tokens you expect to send and receive; cached tokens use the cached-input price where the provider lists one, and count as ordinary input where it does not.

Estimated

—