Skip to content
Sovyron

Llama 3.1 8B Instruct

llama131k context131k max outputreleased Jul 23, 2024Open weightsReasoningTool calling

As of Oct 4, 2026, the cheapest Llama 3.1 8B Instruct API price is $0.02 per 1M input tokens (Kilo Gateway) and $0.025 per 1M output tokens (Inference). Llama 3.1 8B Instruct is offered by 14 providers in the llama family, open weights, reasoning enabled.

Cheapest input
$0.02
Cheapest output
$0.025
Official
—
output per 1M tokens
Providers
14
offering this model
Spread
18.0×
max / min output

API pricing by provider

Metered per-token prices. Cheapest input and output are highlighted; † marks a context-tier surcharge. How prices are ranked.

Llama 3.1 8B Instruct API prices by provider in USD per 1M tokens, updated Oct 4, 2026
Provider1M input1M output1M cache read1M cache writeContext
Inferenceopen$0.025Cheapest$0.025——16k
Kilo GatewayopenCheapest$0.02$0.04——131k
Helicone$0.02$0.05——16k
Abacusopen$0.02$0.05——128k
NovitaAIopen$0.02$0.05——16k
OpenRouteropen$0.05$0.08$0.025—131k
Hugging Faceopen$0.06$0.06——131k
NanoGPTopen$0.0544$0.085$0.0272—131k
Cortecsopen$0.167$0.167——128k
DigitalOceanopen$0.198$0.198——131k
Pioneeropen$0.2$0.2$0.2$0.2128k
Amazon Bedrockopen$0.22$0.22——128k
Amazon Bedrockusopen$0.22$0.22——128k
Vercel AI Gateway$0.22$0.22——128k
Neonopen$0.15$0.45——131k

Llama 3.1 8B Instruct pricing FAQ

What is the cheapest Llama 3.1 8B Instruct API?

As of Oct 4, 2026, Inference has the lowest Llama 3.1 8B Instruct output price at $0.025 per 1M tokens, and Kilo Gateway has the lowest input price at $0.02 per 1M tokens. Prices are metered per-token rates in USD; free tiers and subscription plans are excluded.

How many providers offer Llama 3.1 8B Instruct?

14 providers list Llama 3.1 8B Instruct on Sovyron; 15 of them sell it at a metered per-token price.

Is Llama 3.1 8B Instruct free?

No provider in the Sovyron catalog lists a free tier for Llama 3.1 8B Instruct.

What is the context window of Llama 3.1 8B Instruct?

Llama 3.1 8B Instruct supports a 131k-token context window and up to 131k output tokens.