Evaluate a Voice and Model Without Cherry-Picking the Demo
Compare voice and model choices with a balanced script, blinded listening and task outcomes so one polished sample does not determine the decision.
Table of Contents▼
A voice model evaluation checklist protects against a familiar mistake: choosing the option that sounds best in one short sample. The same voice may struggle with product names, sound tiring in longer explanations or behave differently when the caller interrupts. A useful comparison reflects the work the assistant will do.
Separate the speaking voice from the system's reasoning and action behaviour. If you change the voice, model, prompt and route together, you may prefer the result without knowing what caused the improvement. That makes later troubleshooting difficult.
Build a balanced listening set
Choose phrases from the real workflow, rewritten with fictional data where necessary. Include a greeting, a concise factual answer, an address or code, a clarification and an honest failure response. Add one longer explanation to check whether the style remains comfortable over time.
A hypothetical equipment supplier might test several product names, a serial-number readback and a sentence distinguishing “available to enquire about” from “confirmed in stock.” The goal is intelligibility and appropriate meaning, not theatrical expressiveness.
Burki's TTS comparison can help identify provider questions. It should not replace your own workload-specific listening test or be read as a universal ranking for every language and device.
Keep comparison conditions stable
Use the same text for voice-only comparisons and the same caller objectives for conversational comparisons. Record the voice identifier, model version where exposed, settings, route and test date. If a provider changes its model behind a stable name, your retained samples still provide a historical reference.
Randomise listening order where practical and avoid telling reviewers which sample is expected to win. This reduces brand preference and first-sample effects. A small blinded exercise is useful even when the team has no formal research function.
Let reviewers replay difficult phrases, but preserve their first-listen judgement too. Callers often do not have a replay button during a live conversation.
Score dimensions separately
Use a short scale with written anchors for intelligibility, pronunciation, pacing and tone fit. Add a free-text field for the exact phrase that caused a problem. Do not let “sounds human” stand in for all four dimensions.
Then run a conversational pass for correction handling, interruption recovery and factual answers. LiveKit's simulation documentation describes text and audio evaluation approaches in supported environments. These can assist testing, but human listening remains useful for the specific experience you intend to provide.
A pleasing voice should not compensate for an incorrect answer. Keep task failures separate from preference scores so the final recommendation remains defensible.
Include realistic caller conditions
Listen through the intended telephone or browser route, not only studio headphones. Add a small set of representative devices and acoustic conditions. Note the limitations: a quiet-office headset sample does not describe every mobile caller.
Recruit reviewers familiar with the language and terms being spoken. For multilingual use, evaluate each language and code-switching pattern independently. A provider's supported-language list does not establish that your names, numbers and business vocabulary work well.
Avoid treating one person's accent preference as a quality fact. The useful question is whether the audience understands the assistant and feels comfortable completing the task.
Watch for fatigue over repeated interactions. A style that sounds impressive for ten seconds may become intrusive when every answer uses the same dramatic emphasis. Ask reviewers to complete several ordinary tasks in sequence, then note whether pacing or expressiveness interferes with understanding. This gives the business a more realistic view than repeatedly replaying a favourite greeting.
Choose with tradeoffs visible
Summarise strengths, observed failures and untested conditions. If two options are similar, operational factors such as supported workflow features, account limits and actual cost may decide the choice. Use current account-specific information rather than remembered list prices.
Retest a compact sample when the model, voice or important settings change. Burki's Cartesia and ElevenLabs comparison is a starting point for provider-specific research. The next step is a balanced listening set that includes the awkward phrases your callers will actually need understood.
Ready to try Burki?
Create an assistant and check your available browser practice allowance.
Start Free TrialTrial eligibility and available practice are shown in your workspace.