Voice AI response time: measure the pause callers actually hear
Measure response onset, interruption and tool wait separately, and improve conversational timing without claiming an unverified sub-second guarantee.
Table of Contents▼
A fast model response is not necessarily a fast conversation. The caller experiences the interval between finishing a thought and hearing a useful reply. That interval can include turn detection, transcription, model work, an external lookup and audio playback. Reporting only one component hides the rest.
Burki does not establish a universal sub-second response guarantee through this article. Use the following measurement plan to evaluate the assistant, provider configuration and telephone route you intend to operate.
Define the start and finish of each measurement
For ordinary conversation, mark the last audible part of the caller's completed turn and the first audible part of the assistant's answer. Use a permitted recording or a controlled caller-side observation. A server event saying that audio was generated does not prove when the caller heard it.
For a tool-backed request, record both the acknowledgment and the substantive answer. “Let me check” may arrive quickly while the customer still waits for a booking result. Report that distinction instead of counting filler as completed service.
| Measure | Start | Finish |
|---|---|---|
| Response onset | Caller completes their turn | Assistant becomes audible |
| Interruption response | Caller begins a genuine interruption | Assistant playback stops |
| Action completion | Confirmed request is submitted | Verified result reaches the caller |
| Recovery delay | Failure becomes known | Caller receives a useful next step |
Test pauses as well as complete sentences
A caller might say, “My address is… let me check… fourteen Oak Street.” A detector that responds to the first pause can feel quick while repeatedly cutting people off. Include hesitations, corrections, brief acknowledgments and background speech in the test set.
LiveKit's turn-handling documentation distinguishes turn detection from interruption handling and describes several detection strategies. Those controls affect when a response starts; their names do not establish that your configuration is well tuned.
A useful correction test is: “Actually, not Tuesday—Thursday.” Check whether the assistant stops, retains the correction and asks the next relevant question. A short stop time alone is insufficient if the next answer still uses Tuesday.
Compare like with like
Keep the prompt, business facts, caller script and tools stable when comparing configurations. Record the region, connection type, model, voice and assistant revision. Test browser and telephone routes separately because they exercise different paths.
Report the sample size, median and a high percentile rather than only the fastest exchange. With a small sample, show individual slow turns as well; one percentile can imply more certainty than the experiment supports. Categorize delays by ordinary answers, external actions and recovery.
Change the part responsible for the delay
If a lookup is slow, inspect that dependency before replacing the voice. If the assistant waits after every short answer, review turn handling. If it gives long speeches, shorten the instructions and ask one question at a time. Recheck accuracy after each change.
Burki's browser practice provides a place to rehearse these scenarios before connecting callers. Keep a separate acceptance record for live actions and telephone behavior. The goal is a conversation that listens and resolves the request, supported by measured timing rather than a latency slogan.
Ready to try Burki?
Create an assistant and check your available browser practice allowance.
Create your assistantTrial eligibility and available practice are shown in your workspace.