Skip to content
Sovyron

Inference

9 modelsProvider docs

As of Oct 4, 2026, Inference lists 9 supported models in the Sovyron catalog, including Google Gemma 3, Llama 3.1 8B Instruct, Llama 3.2 11B Vision Instruct, Llama 3.2 1B Instruct and Llama 3.2 3B Instruct — gemma, llama, mistral-nemo, osmosis and qwen families. Its metered rates span $0.0075–$0.2 per 1M tokens on a 3:1 input:output blend, and every price in the table is Inference's own, normalized to USD — open a model to compare it with the cheapest offer across all providers.

Supported models and pricing

9 of 9 models
Inference model prices in USD per 1M tokens, updated Oct 4, 2026
Model1M input1M output1M cache readContextReleased
Google Gemma 3$0.15$0.3—125kJan 1, 2025
Llama 3.1 8B Instruct$0.025$0.025—16kJan 1, 2025
Llama 3.2 11B Vision Instruct$0.055$0.055—16kJan 1, 2025
Llama 3.2 1B Instruct$0.01$0.01—16kJan 1, 2025
Llama 3.2 3B Instruct$0.02$0.02—16kJan 1, 2025
Mistral Nemo 12B Instruct$0.038$0.1—16kJan 1, 2025
Osmosis Structure 0.6B$0.1$0.5—4kJan 1, 2025
Qwen 2.5 7B Vision Instruct$0.2$0.2—125kJan 1, 2025
Qwen 3 Embedding 4B$0.01free—32kJan 1, 2025