The reference shape is OpenAI’s REST API: a POST to https://api.openai.com/v1/chat/completions with an Authorization: Bearer header (OpenAI API reference, read 6 October 2026). Compatibility layers such as Anthropic’s and OVHcloud’s implement that Chat Completions format. OpenAI itself keeps Chat Completions supported but recommends its newer Responses API for all new projects (migration guide, read 6 October 2026), so check which of the two a provider means.
What does OpenAI compatibility cover?
It covers endpoint paths, bearer authentication, the messages array, model names as plain strings, and the response fields an SDK reads. That is enough to point an existing client at another inference provider without new adapter code. Streaming events, tool calling, structured output, multimodal inputs and error bodies are where implementations drift.
API compatibility is not model compatibility.
Where does OpenAI compatibility break?
Anthropic documents its own layer in detail. It ignores response_format and the strict flag on tools, so tool-call JSON is not guaranteed to follow your schema, and it merges every system message into one at the start of the conversation (OpenAI SDK compatibility, read 6 October 2026). It says unsupported fields mostly fail silently instead of returning errors, and positions the layer as “primarily intended to test and compare model capabilities.” A silently ignored field is harder to catch than a 400 error, because nothing fails. Compatibility also runs the other way: OpenAI’s Codex CLI accepts custom model providers, but responses is the only supported wire_api value (Codex config reference, read 6 October 2026), so a provider that speaks only Chat Completions cannot back it.
How do we use an OpenAI-compatible API at Soba Labs?
The LLM judge that reviews our pull requests calls the Chat Completions endpoint of OVHcloud AI Endpoints with plain HTTP from Python’s standard library. There is no SDK: one POST with a bearer token and a JSON body, and the model is a single configuration value. Because every model on that endpoint takes the same request, choosing among them was a scoring question, not integration work. The scores behind our pick, checked on 23 September 2026: Qwen3.8-27B at 34 on the Artificial Analysis index v4.3.2, against 18 for Qwen3.5-397B-A17B and 12 for gpt-oss-120b.
The shared format still left gaps, covered on the inference provider page: a rejected vLLM parameter and reasoning returned in a separate field. So the parser never reads an empty answer as a verdict: it raises an error recording the finish reason and token usage, and it removes any inline think block before reading the JSON. The request also names its reasoning effort instead of trusting the provider’s default.
How should teams test an OpenAI-compatible provider?
Treat compatibility as a transport convenience and contract-test each provider for the features you rely on: plain output, streaming, tools, structured output, limits and error bodies. Assert the effect of every parameter you send, not just a 200 response. Keep provider-specific exceptions in one adapter or an LLM gateway, which keeps model portability honest, and record the provider and model of every call in LLM observability so a failure traces to the route that served it.