Inference Providers: What They Are and How We Choose One

ProductionModels and inferencePublished Updated By Simon Budziak

An inference provider is a service that runs trained AI models and returns their outputs over an API. It owns the accelerators, scaling, batching and runtime tuning, so your team sends requests instead of operating GPUs. The model can come from the provider, from another lab, or from the customer's own open weights.

Providers such as Groq, Together AI and OVHcloud AI Endpoints sell this serving layer. All three appear among the partners behind Hugging Face’s Inference Providers router (read 6 October 2026), which by default sends each request to the fastest available provider for the chosen model and switches to the lowest price per output token when the model name ends in :cheapest.

What does an inference provider manage?

The provider turns a model file into a reachable production service. It handles most of the model serving work an application team would otherwise run itself: model loading, hardware, autoscaling, queues, logging and regional endpoints. Most expose an OpenAI-compatible API, so moving between providers is often a change of base URL and model name.

The same model can behave differently across providers because serving configuration changes latency, limits and accepted parameters.

Inference provider vs LLM provider: what’s the difference?

An LLM provider is anyone who gives you access to a language model, including the labs that train them. An inference provider is defined by the serving job, not by who built the model. Many inference providers mostly serve open-weight models trained by someone else, so the provider rather than the model’s author becomes your contractual counterparty, a point we worked through in Can you use a Chinese AI model?.

What have we learned running an inference provider at Soba Labs?

Two of our own production jobs run on OVHcloud AI Endpoints, which says its LLM APIs follow the OpenAI specification and that it does not store user data (getting started guide, updated 19 March 2026): an LLM judge on Qwen3.8-27B that reviews every pull request in our own repositories, and speech-to-text for our video pipeline on whisper-large-v3.

We chose the judge model on 23 September 2026 because it had the highest Artificial Analysis index score among the models that provider served. Three things we hit while wiring it up now shape every call:

So the judge gets one call at temperature 0, after cheap deterministic checks that stop a failing pull request before it reaches the model. It retries once on a server or network error and never on a 4xx, and any answer outside its JSON contract goes to a human instead of merging or rejecting.

How should inference providers be compared?

Pin the same model and settings, then send production-shaped prompts at expected concurrency. Measure queue time, time to first token, total latency, errors, cold starts and cost per accepted output. Review region, data handling and model-update policy separately from speed. The lowest token price can be the most expensive route if retries or slow answers cut completed work. Log the provider and outcome of every request with LLM observability, especially behind a model router.

Written with AI assistance and reviewed by Simon Budziak. The production notes come from systems Soba Labs builds and runs.

Frequently asked questions

What are inference providers?

Inference providers are services that host trained AI models and return their outputs over an API, so a team can call a model without buying or operating its own accelerators.

Does an inference provider create the model?

Not necessarily. Many serve open-weight models from other labs, and some also host models a customer trained or fine-tuned.

Is an inference provider the same as an LLM provider?

Not always. LLM provider covers anyone giving access to a language model, including the labs that train them, while an inference provider is defined by running models, often ones built elsewhere.

Summarize this page with

Train your team to build this