What are the main types of LLM provider?
Four kinds of company answer to the name:
- Model developers such as OpenAI and Anthropic train models and sell access through their own API.
- Cloud platforms such as Amazon Bedrock offer several labs’ models inside your existing cloud account and contract.
- Inference providers host open-weight models that others trained.
- Routers such as OpenRouter put many providers behind one endpoint, and our analysis of its usage data shows how much traffic now flows through that layer.
The type decides who your contractual counterparty is, and that matters more for data terms than the model’s name.
What does an LLM provider manage, and what stays with you?
The provider handles model hosting, inference, API access, capacity and model updates. You still own the application logic, permissions, evaluation and the handling of model failures. A model API is a dependency inside a system, not the system itself. Compare candidates on your own representative tasks: task success, structured output reliability, latency, rate limits, regions, data terms and cost per accepted output, weighted before results arrive and kept in an AI evaluation harness. Review AI vendor lock-in before provider-specific features spread through the codebase.
Which LLM providers does Soba Labs use, and why more than one?
We pick a provider per job, not one for everything. Our coding agents run on Anthropic and OpenAI models, and on delivery work the review goes to the provider that did not write the code: work built with Claude Code gets a review from OpenAI’s Codex, and work built with Codex gets a Claude review. Since 20 July 2026 the Codex reviews run on OpenAI’s GPT-5.6 family at high reasoning effort. We trialed the highest effort on 6 July 2026 and reverted it as too slow for our review cadence, because a high-effort review of a real diff already takes 15 to 30 minutes or more. Our pull request judge and our transcription run on models hosted by an inference provider, described on the inference provider page. Running several providers on purpose is what multi-provider AI looks like in practice.
What have we learned about getting access to a provider’s models?
Seeing a model in a cloud console does not mean your application can call it. When we build inside a client’s AWS account, access splits into separate checks: the runtime permission to invoke the model, the region the model actually runs in, model enablement, and the secrets, logging and network permissions the application needs. The only definitive test is to invoke the model from the application’s own role and region. AWS documents why: the first call to a third-party model starts a background subscription that can take up to 15 minutes, Anthropic models on the standard endpoint need a one-time use case form first, and some access is decided separately for each account (Bedrock model access, read 6 October 2026).