What latency budget does a voice agent need?
A voice agent’s budget is tight because people expect a reply almost as soon as they stop talking. In a study of 10 languages, answers to yes or no questions started on average 208 milliseconds after the question ended (Stivers et al., PNAS, 2009, read 6 October 2026). ITU-T G.114 caps one-way telephony delay at 400 ms for network planning and notes that many voice calls are affected at much lower delays (ITU-T G.114 summary, in force, read 6 October 2026).
Endpointing can eat most of the budget, and it is a setting you choose. LiveKit’s agent framework waits at least 0.5 seconds after the last detected speech before ending the turn by default (LiveKit turn handling options, updated 6 October 2026). The diagram’s example gives endpointing 140 ms of 800 ms, so that default alone spends more than half the budget. Review turn detection and barge-in settings together with the budget.
How should a latency budget be set and measured?
Start where waiting becomes disruptive for the task, then divide that target among the stages that consume it. Every component inherits a limit from the experience it serves. Measure the event the user notices: time to first token for text, the first audible word for voice, the confirmed result for a tool. Report p95 beside the median, keep timeout and failure counts next to latency, and timestamp every boundary through LLM observability, because a component can meet its local target while the whole interaction misses.
How do we set a latency budget for an agent?
We budget the worst case, which is timeouts multiplied by retries. Our open-source KRS MCP server, an agent tool for Polish company records, gives each upstream request 15 seconds and retries network errors, HTTP 429 and server errors twice, after 250 ms and 750 ms. One tool call can take about 46 seconds before the agent learns it failed. The timer also stays armed while the response body downloads, since clearing it when headers arrive leaves a stalled body unbounded.
Order matters too. In one agent we operated, a cheap check sat behind two throttled lookups of about 50 seconds each: instant in development, past its 120 second tool deadline under load. Now the cheap deterministic check runs first. In LangGraph the node timeout resets on each retry, so the retry policy multiplies the agent timeout. In our own pipelines the model call is the big line item: in the LLM review step we run on every pull request, a healthy run takes about a minute, and 25 to 45 seconds of it is the judge model.
What happens when the budget is missed?
Decide the fallback before production: a faster model, a skipped enrichment step, a partial result, or a handoff to a person. A latency fallback must preserve correctness and permissions. On a call, acknowledge quickly and do the slow work after, as in voice agents in production, but never promise what a tool has not confirmed. Optimize the stage that consumes the tail, not the stage easiest to rewrite, then re-measure.