Back to Blog
BYO & Cost Control

ElevenLabs vs Deepgram vs Cartesia: Choosing TTS for Voice Agents

Compare speech models, streaming behavior, pronunciation and complete call costs, then use a repeatable listening test to choose a voice.

Burki
(Updated: September 25, 2026)
4 min read

The best text-to-speech provider for a voice agent is the one that callers can understand, interrupt and finish a task with. A convincing sample recording is a useful first filter. It does not show how a voice behaves when a caller changes an address, asks a follow-up question or speaks over a long answer.

This guide compares the decisions involved in choosing ElevenLabs, Deepgram and Cartesia. Provider documentation was reviewed on September 25, 2026. The model facts below come from those providers; we have not run a controlled three-provider benchmark and do not claim a universal speed, quality or price winner.

Compare models, not just company names

ProviderWhat its current catalog distinguishesWhat to put on your shortlist
ElevenLabsFlash, Multilingual and expressive v3 families have different intended usesAn exact model and voice that support your language and conversational needs
DeepgramFlux TTS is separate from Aura-2; the published language coverage differsThe appropriate TTS family for your language, not a similarly named transcription model
CartesiaSonic provides streaming speech; its current documentation identifies Sonic 3.6 and dated snapshotsA specific Sonic version and voice, with an upgrade policy

Check the current ElevenLabs model catalog, Deepgram TTS overview and Cartesia model documentation before choosing. A provider releasing a model does not mean every voice platform has integrated that version.

First decide whether you need a separate TTS provider

In a chained voice assistant, a transcription service produces text, a language model decides what to say, and TTS turns the answer into audio. A speech-to-speech architecture can handle audio directly. Those are different purchasing decisions; swapping the TTS box is not how every realtime system changes its voice.

If you are comparing a complete voice agent with a standalone synthesis API, start with the OpenAI Realtime and ElevenLabs architecture guide. If you are choosing the component that hears callers, use the Deepgram versus ElevenLabs guide.

Use a listening test that resembles your calls

Prepare the same short script for each exact voice and model. Include a greeting, a street address, an appointment time with a timezone, a product code, an unfamiliar surname and one sentence in each required language. Label these as test examples; use no customer secrets.

Then test inside a conversation. Interrupt during a long answer and ask for a correction. Record whether the old audio actually stops, whether the corrected detail is repeated accurately, and whether the next response remains concise. Compare the audio callers hear through the intended telephone route as well as browser audio.

ObservationHow to record it
IntelligibilityWhich names, numbers or words listeners misunderstood
Response onsetTime from a completed caller turn to audible assistant speech
InterruptionWhether obsolete audio continues after the caller starts correcting it
ConsistencyWhether the voice or pronunciation changes across turns
Task completionWhether the caller's requested outcome was correctly completed

Keep provider synthesis time separate from complete response time. Transcription, turn detection, model reasoning, tools, buffering and the network can all contribute to the pause a caller experiences. A vendor's first-audio figure is not a measurement of your whole call.

Compare cost using the same workload

Request a quote for the selected model, voice, language and expected usage. Note the billing unit, any minimum commitment, included allowance, concurrency limits and optional voice charges. Measure actual generated speech during a representative conversation; a verbose assistant may generate more audio than your budget assumed.

Add transcription, reasoning, the voice platform, telephone minutes and number rental where applicable. Do not compare one provider's synthesis component with another platform's complete call price. Burki's pricing page explains its own usage categories; provider-direct bills may remain separate when using your credentials.

Choose, then verify the complete assistant

Shortlist two voices that pass your language and pronunciation checks. Choose between them using the task score and the full cost, then repeat the same checks before changing model versions.

In Burki, begin with a business receptionist draft and the voice modes actually available to your workspace. Test the conversation before making a provider choice the centerpiece of the project. A caller needs an accurate answer and a reliable next step more than a long list of voice options.

Ready to try Burki?

Create an assistant and check your available browser practice allowance.

Create your assistant

Trial eligibility and available practice are shown in your workspace.

Related Articles