As of Oct 4, 2026, Inference lists 9 supported models in the Sovyron catalog, including Google Gemma 3, Llama 3.1 8B Instruct, Llama 3.2 11B Vision Instruct, Llama 3.2 1B Instruct and Llama 3.2 3B Instruct — gemma, llama, mistral-nemo, osmosis and qwen families. Its metered rates span $0.0075–$0.2 per 1M tokens on a 3:1 input:output blend, and every price in the table is Inference's own, normalized to USD — open a model to compare it with the cheapest offer across all providers.
Supported models and pricing
9 of 9 models
Inference model prices in USD per 1M tokens, updated Oct 4, 2026