Skip to content
Sovyron

Llama 3.3 70B

llama131k context131k max outputreleased Dec 6, 2024Open weightsTool calling

As of Oct 4, 2026, the cheapest Llama 3.3 70B API price is $0.53 per 1M input tokens (STACKIT) and $0.71 per 1M output tokens (CoreWeave). Llama 3.3 70B is offered by 6 providers in the llama family, open weights.

Cheapest input
$0.53
Cheapest output
$0.71
Official
—
output per 1M tokens
Providers
6
offering this model
Spread
3.9×
max / min output
Free offers
0
free-tier offers

API pricing by provider

Metered per-token prices. Cheapest input and output are highlighted; † marks a context-tier surcharge. How prices are ranked.

Llama 3.3 70B API prices by provider in USD per 1M tokens, updated Oct 4, 2026
STACKITopenCheapest$0.53$0.76——128k
Groqopen$0.59$0.79——131k
CoreWeaveopen$0.71Cheapest$0.71$0.71—128k
Together AIopen$1.04$1.04——131k
Venice AIopen$0.7$2.8——128k
NanoGPTopen$1.75$2.75$1.75—128k

Llama 3.3 70B pricing FAQ

What is the cheapest Llama 3.3 70B API?

As of Oct 4, 2026, CoreWeave has the lowest Llama 3.3 70B output price at $0.71 per 1M tokens, and STACKIT has the lowest input price at $0.53 per 1M tokens. Prices are metered per-token rates in USD; free tiers and subscription plans are excluded.

How many providers offer Llama 3.3 70B?

6 providers list Llama 3.3 70B on Sovyron; 6 of them sell it at a metered per-token price.

Is Llama 3.3 70B free?

No provider in the Sovyron catalog lists a free tier for Llama 3.3 70B.

What is the context window of Llama 3.3 70B?

Llama 3.3 70B supports a 131k-token context window and up to 131k output tokens.

Token cost calculator

What a month costs at one provider's published rates. Enter the tokens you expect to send and receive; cached tokens use the cached-input price where the provider lists one, and count as ordinary input where it does not.

Estimated

—