Back to Blog
Provider Integrations

Deepgram for voice agents: Nova, Flux and the turn-taking boundary

Choose and evaluate a Deepgram transcription path using language, vocabulary, turn events, caller corrections and actual call evidence.

Burki
(Updated: September 25, 2026)
3 min read

Speech recognition quality affects more than the transcript. It influences whether a voice assistant captures names correctly, waits for the caller to finish and responds to a correction. Evaluate the transcription model and the turn-handling configuration together.

Deepgram documents multiple model and language options, including Nova and Flux. Consult the current model overview and Flux quickstart for the intended API and behavior. Do not assume their options or event semantics are interchangeable.

Identify the problem first

Observed problemEvidence to inspect
Names or addresses are wrongRepresentative audio and the recognized words
Assistant responds too earlyAudio timing and end-of-turn events
Assistant waits too longEndpointing, model output and playback timeline
Corrections are ignoredRecognition result and conversation state
Noise triggers responsesAudio conditions and interruption policy

A transcript error can originate in poor audio or an unsupported language setting. A slow answer can occur after recognition is already complete. Do not change the transcription provider until you have located the relevant stage.

Configure the selected model's supported options

Burki's LiveKit cascade runtime has separate Deepgram Nova and Flux paths. Model-specific validation matters: an option supported by one path may not apply to the other.

Use only the vocabulary, language and turn settings exposed for that selected configuration. Supplying a long business prompt is not a substitute for a supported recognition-context mechanism, and contextual hints do not justify inventing words absent from the audio.

Build a small recognition test set

Include your business name, common services, local street names, phone numbers and a caller changing a previously stated detail. Test ordinary phone audio as well as the quiet browser microphone you use during setup.

For each item, record the intended words, recognized words and whether the assistant's next action was correct. Not every transcription difference changes the outcome: punctuation may be harmless while one incorrect digit in a callback number is critical.

Observe turn handling in real conversation

Ask a question with a pause in the middle. Add a short acknowledgment while the assistant speaks, then interrupt with a substantive correction. Check whether the system distinguishes those situations well enough for the task.

A detector or model name is not evidence of good interruption behavior. Measure audible stopping and response timing with representative calls, then compare configurations under the same conditions.

Reconcile quality and cost

Include transcription usage, model work, generated speech and telephone charges where applicable. Repeated requests caused by recognition errors can increase the total even if the transcription unit price is low.

Use the assistant's readiness checks and current pricing before funded tests. For troubleshooting across the whole voice path, see the phone-call API guide. The goal is correct caller information and sensible timing, not a benchmark percentage copied from unrelated audio.

Ready to try Burki?

Create an assistant and check your available browser practice allowance.

Create your assistant

Trial eligibility and available practice are shown in your workspace.

Related Articles