ElevenLabs vs Deepgram vs Cartesia: Choosing TTS for Voice Agents
Compare speech models, streaming behavior, pronunciation and complete call costs, then use a repeatable listening test to choose a voice.
Table of Contents▼
The best text-to-speech provider for a voice agent is the one that callers can understand, interrupt and finish a task with. A convincing sample recording is a useful first filter. It does not show how a voice behaves when a caller changes an address, asks a follow-up question or speaks over a long answer.
This guide compares the decisions involved in choosing ElevenLabs, Deepgram and Cartesia. Provider documentation was reviewed on September 25, 2026. The model facts below come from those providers; we have not run a controlled three-provider benchmark and do not claim a universal speed, quality or price winner.
Compare models, not just company names
| Provider | What its current catalog distinguishes | What to put on your shortlist |
|---|---|---|
| ElevenLabs | Flash, Multilingual and expressive v3 families have different intended uses | An exact model and voice that support your language and conversational needs |
| Deepgram | Flux TTS is separate from Aura-2; the published language coverage differs | The appropriate TTS family for your language, not a similarly named transcription model |
| Cartesia | Sonic provides streaming speech; its current documentation identifies Sonic 3.6 and dated snapshots | A specific Sonic version and voice, with an upgrade policy |
Check the current ElevenLabs model catalog, Deepgram TTS overview and Cartesia model documentation before choosing. A provider releasing a model does not mean every voice platform has integrated that version.
First decide whether you need a separate TTS provider
In a chained voice assistant, a transcription service produces text, a language model decides what to say, and TTS turns the answer into audio. A speech-to-speech architecture can handle audio directly. Those are different purchasing decisions; swapping the TTS box is not how every realtime system changes its voice.
If you are comparing a complete voice agent with a standalone synthesis API, start with the OpenAI Realtime and ElevenLabs architecture guide. If you are choosing the component that hears callers, use the Deepgram versus ElevenLabs guide.
Use a listening test that resembles your calls
Prepare the same short script for each exact voice and model. Include a greeting, a street address, an appointment time with a timezone, a product code, an unfamiliar surname and one sentence in each required language. Label these as test examples; use no customer secrets.
Then test inside a conversation. Interrupt during a long answer and ask for a correction. Record whether the old audio actually stops, whether the corrected detail is repeated accurately, and whether the next response remains concise. Compare the audio callers hear through the intended telephone route as well as browser audio.
| Observation | How to record it |
|---|---|
| Intelligibility | Which names, numbers or words listeners misunderstood |
| Response onset | Time from a completed caller turn to audible assistant speech |
| Interruption | Whether obsolete audio continues after the caller starts correcting it |
| Consistency | Whether the voice or pronunciation changes across turns |
| Task completion | Whether the caller's requested outcome was correctly completed |
Keep provider synthesis time separate from complete response time. Transcription, turn detection, model reasoning, tools, buffering and the network can all contribute to the pause a caller experiences. A vendor's first-audio figure is not a measurement of your whole call.
Compare cost using the same workload
Request a quote for the selected model, voice, language and expected usage. Note the billing unit, any minimum commitment, included allowance, concurrency limits and optional voice charges. Measure actual generated speech during a representative conversation; a verbose assistant may generate more audio than your budget assumed.
Add transcription, reasoning, the voice platform, telephone minutes and number rental where applicable. Do not compare one provider's synthesis component with another platform's complete call price. Burki's pricing page explains its own usage categories; provider-direct bills may remain separate when using your credentials.
Choose, then verify the complete assistant
Shortlist two voices that pass your language and pronunciation checks. Choose between them using the task score and the full cost, then repeat the same checks before changing model versions.
In Burki, begin with a business receptionist draft and the voice modes actually available to your workspace. Test the conversation before making a provider choice the centerpiece of the project. A caller needs an accurate answer and a reliable next step more than a long list of voice options.
Ready to try Burki?
Create an assistant and check your available browser practice allowance.
Create your assistantTrial eligibility and available practice are shown in your workspace.