LLM Provider Meaning: Model Labs, Clouds and Inference Hosts

ProductionModels and inferencePublished Updated By Simon Budziak

An LLM provider is a company or platform that gives your software access to large language models through an API. Some train the models, like OpenAI or Anthropic. Some resell several labs' models on their cloud, like Amazon Bedrock. Some only host open weights. The choice sets quality, cost, latency, data terms and switching effort.

What are the main types of LLM provider?

Four kinds of company answer to the name:

The type decides who your contractual counterparty is, and that matters more for data terms than the model’s name.

What does an LLM provider manage, and what stays with you?

The provider handles model hosting, inference, API access, capacity and model updates. You still own the application logic, permissions, evaluation and the handling of model failures. A model API is a dependency inside a system, not the system itself. Compare candidates on your own representative tasks: task success, structured output reliability, latency, rate limits, regions, data terms and cost per accepted output, weighted before results arrive and kept in an AI evaluation harness. Review AI vendor lock-in before provider-specific features spread through the codebase.

Which LLM providers does Soba Labs use, and why more than one?

We pick a provider per job, not one for everything. Our coding agents run on Anthropic and OpenAI models, and on delivery work the review goes to the provider that did not write the code: work built with Claude Code gets a review from OpenAI’s Codex, and work built with Codex gets a Claude review. Since 20 July 2026 the Codex reviews run on OpenAI’s GPT-5.6 family at high reasoning effort. We trialed the highest effort on 6 July 2026 and reverted it as too slow for our review cadence, because a high-effort review of a real diff already takes 15 to 30 minutes or more. Our pull request judge and our transcription run on models hosted by an inference provider, described on the inference provider page. Running several providers on purpose is what multi-provider AI looks like in practice.

What have we learned about getting access to a provider’s models?

Seeing a model in a cloud console does not mean your application can call it. When we build inside a client’s AWS account, access splits into separate checks: the runtime permission to invoke the model, the region the model actually runs in, model enablement, and the secrets, logging and network permissions the application needs. The only definitive test is to invoke the model from the application’s own role and region. AWS documents why: the first call to a third-party model starts a background subscription that can take up to 15 minutes, Anthropic models on the standard endpoint need a one-time use case form first, and some access is decided separately for each account (Bedrock model access, read 6 October 2026).

Written with AI assistance and reviewed by Simon Budziak. The production notes come from systems Soba Labs builds and runs.

Frequently asked questions

What does LLM provider mean?

It means any company or platform that gives software access to large language models, whether it trains the models, resells them on a cloud, or only hosts open weights.

What is the difference between an LLM provider and an inference provider?

An LLM provider may develop or distribute models. An inference provider focuses on running models and returning outputs, including models created elsewhere.

How many LLM providers should an application use?

Use the fewest that meet a real requirement. Add another for a measured capability, resilience, regional or cost need, or so one provider's model can review another's work.

Summarize this page with

Train your team to build this