Providers such as Groq, Together AI and OVHcloud AI Endpoints sell this serving layer. All three appear among the partners behind Hugging Face’s Inference Providers router (read 6 October 2026), which by default sends each request to the fastest available provider for the chosen model and switches to the lowest price per output token when the model name ends in :cheapest.
What does an inference provider manage?
The provider turns a model file into a reachable production service. It handles most of the model serving work an application team would otherwise run itself: model loading, hardware, autoscaling, queues, logging and regional endpoints. Most expose an OpenAI-compatible API, so moving between providers is often a change of base URL and model name.
The same model can behave differently across providers because serving configuration changes latency, limits and accepted parameters.
Inference provider vs LLM provider: what’s the difference?
An LLM provider is anyone who gives you access to a language model, including the labs that train them. An inference provider is defined by the serving job, not by who built the model. Many inference providers mostly serve open-weight models trained by someone else, so the provider rather than the model’s author becomes your contractual counterparty, a point we worked through in Can you use a Chinese AI model?.
What have we learned running an inference provider at Soba Labs?
Two of our own production jobs run on OVHcloud AI Endpoints, which says its LLM APIs follow the OpenAI specification and that it does not store user data (getting started guide, updated 19 March 2026): an LLM judge on Qwen3.8-27B that reviews every pull request in our own repositories, and speech-to-text for our video pipeline on whisper-large-v3.
We chose the judge model on 23 September 2026 because it had the highest Artificial Analysis index score among the models that provider served. Three things we hit while wiring it up now shape every call:
- Reasoning effort is a setting you have to test per provider. At the provider’s default, which is also its highest level, the model spent 16,000 tokens thinking about a 72 KB diff and never answered. At medium it used about 4,500 reasoning tokens and answered in 44 seconds.
- The provider rejected, with a 400 error, a request parameter that vLLM, the open-source serving engine, accepts. “OpenAI compatible” does not mean every option passes through.
- The reasoning comes back in a separate field, and the answer field is empty if the token budget runs out first.
So the judge gets one call at temperature 0, after cheap deterministic checks that stop a failing pull request before it reaches the model. It retries once on a server or network error and never on a 4xx, and any answer outside its JSON contract goes to a human instead of merging or rejecting.
How should inference providers be compared?
Pin the same model and settings, then send production-shaped prompts at expected concurrency. Measure queue time, time to first token, total latency, errors, cold starts and cost per accepted output. Review region, data handling and model-update policy separately from speed. The lowest token price can be the most expensive route if retries or slow answers cut completed work. Log the provider and outcome of every request with LLM observability, especially behind a model router.