Back to Blog
Provider Integrations

ElevenLabs voice integration: choose a model and test the whole call

Configure a supported ElevenLabs voice for a cascade assistant and evaluate pronunciation, streaming, interruption and generated-speech cost.

Burki
(Updated: September 25, 2026)
3 min read

An ElevenLabs voice can be one component of a voice assistant, but a convincing preview does not establish how it behaves during a live call. The model, voice, language, streaming path and interruption handling all affect the result.

Burki's cascade runtime includes an ElevenLabs synthesis adapter. Availability depends on supported configuration, credentials and pricing readiness. This does not mean an ElevenLabs voice can simply replace the native speech output of every GPT Live configuration.

Select the voice and model together

ElevenLabs maintains a model catalog with different capabilities and intended uses. Consult the current model documentation rather than choosing from an old model-name or latency table.

Verify that the chosen voice is available to your account and appropriate for the language and task. For cloned or shared voices, confirm you have the necessary rights and provider permissions before using them with callers.

Prepare a realistic voice sample

Use the actual words your assistant will say: business name, street names, phone numbers, appointment times and common service terms. Include a short greeting and a longer explanation.

Listen for intelligibility, pacing and pronunciation. An expressive promotional sample may sound different from a concise appointment confirmation over a telephone connection.

Test interaction rather than playback alone

ScenarioWhat to observe
Caller interrupts the greetingSpeech stops promptly and the new request is handled
Caller corrects a detailThe next response uses the corrected information
Tool is slowThe assistant does not fill the wait with excessive speech
Action failsThe voice communicates an honest next step
Call ends mid-responseResources stop and usage settles coherently

Measure what the caller hears. A transcript showing the full response does not prove that all generated audio was played, and a stopped playback does not prove no provider work was already consumed.

Control generated-speech expense

Keep responses proportional to the task. Ask one question at a time, avoid repeating the entire request after every answer, and confirm only the details that matter. Compare actual provider usage rather than estimating solely from connected minutes.

Managed synthesis requires an applicable rate and usage accounting; a provider option without pricing readiness should not be treated as free. With BYO credentials, include the direct provider invoice alongside platform and carrier expense.

Verify before publishing

Save and test the selected draft with representative phrases, then run a permitted telephone test if real callers will use it. Keep the previous configuration available until the replacement passes your scenarios.

Read BYO versus managed usage for account responsibilities and cost tracking for reconciliation. Choose the voice that makes the business conversation clearer, not simply the most impressive isolated audio sample.

Ready to try Burki?

Create an assistant and check your available browser practice allowance.

Create your assistant

Trial eligibility and available practice are shown in your workspace.

Related Articles