What values does reasoning effort accept?
OpenAI takes reasoning.effort in the Responses API and reasoning_effort in Chat Completions, with model-dependent values that can include none, minimal, low, medium, high, xhigh and max (OpenAI reasoning guide, read 6 October 2026); GPT-6 Astra, for one, returns HTTP 400 for none. Anthropic takes output_config.effort with low, medium, high, xhigh or max; most Claude models default to high and Claude Opus 5.5 to medium (Anthropic effort docs, read 6 October 2026). Anthropic treats effort as a steer on behavior, not a hard token limit. It is also a separate control from temperature on a reasoning model.
Why is reasoning effort a cost and latency decision?
Every level up spends more thinking tokens on every call. OpenAI warns that a response can hit max_output_tokens before any visible output, billing input and reasoning with no answer, and suggests reserving at least 25,000 tokens when you start. Effort buys accuracy on hard problems and little on easy ones. Set it per task, not per application, measure whether a higher level changes the outcome, and let model routing choose model and effort together. Where a user is waiting, your latency budget caps it, and the total belongs in AI FinOps reporting.
How do we set and verify reasoning effort at Soba Labs?
We name the effort on every request instead of trusting a default. The LLM judge that reviews our pull requests runs at medium, because the provider’s default, its highest level, never answered on a large diff (the full story is on the inference provider page). The same call carries a 10,000 token output cap and a 400 second HTTP timeout, and the job around it stops after 10 minutes.
Codex taught us to check what ran, not what we sent. An unquoted effort override on the Codex command line landed a run at medium with no warning, while the quoted TOML string form worked. Codex’s advanced config page (read 6 October 2026) now says a value that does not parse as TOML is treated as a string, so behavior can shift between releases (see our Codex CLI post). The Codex plugin we run from Claude Code took a model flag but no effort flag on its review command in version 1.0.5 (checked 23 August 2026), so we force the high effort our reviews use, described on the LLM provider page, through its general task command. Either way we confirm model and effort from the session log Codex writes for each turn.
Low effort changes prompting too: a low-effort model does less interpretive repair, so we write API versions, enum values and IDs into its prompt literally rather than let it invent a plausible default.
Where do reasoning effort settings fail silently?
Anthropic’s OpenAI SDK compatibility layer is one documented case: it lists reasoning_effort as ignored and says unsupported fields mostly fail silently instead of returning errors (OpenAI SDK compatibility, read 6 October 2026). A request shaped for an OpenAI-compatible API can therefore succeed while the effort you chose does nothing. Treat effort as unset until a log, a token count or the provider’s own record shows it applied.