VOICE-AGENT STT COMPARISON WORKSHEET Free planning resource. Not a provider benchmark or a promise of free API usage. Prepared October 6, 2026. 1. Define the decision Workflow: Target caller languages and accents: Actual browser/telephone route: Critical fields that must be correct: Approved test budget / authorization (leave blank until authorized): 2. Freeze the comparison Audio-set identifier and consent/provenance: Human-checked reference transcript and expected field values: Provider A / exact model / alias or immutable snapshot: Provider B / exact model / alias or immutable snapshot: Account region / request options / sample rate / encoding: Prompt/keyterm settings (record a baseline before tuning): SDK/runtime versions: Evaluation date: Same replay order and equivalent test conditions: 3. Use fictional examples first CASE A: Caller says 'I need a callback, not an appointment.' Expected: callback request; no booked appointment claim. CASE B: Caller says 'My name is Wexley. W E X L E Y.' Expected: retain corrected spelling. CASE C: Caller provides a fictional phone number, pauses, then corrects one digit. Expected: turn stays open or clarification preserves final confirmed value. CASE D: Caller says only 'no' after a confirmation question. Expected: refusal, no execution of the declined action. CASE E: Same request in each supported target language and an approved code switch. Expected: accurate intent/fields; unsupported language handled honestly. CASE F: Consented version with realistic background noise. Expected: do not invent words from another speaker. 4. Record observations, not just a leaderboard Case ID: Final reference / final transcript: Word substitutions / deletions / insertions / reference word count: WER = (substitutions + deletions + insertions) / reference words. Use the same normalization; do not mix semantic WER with ordinary WER. Critical field value correct? yes/no/uncertain: Caller end of speech timestamp: Final transcript timestamp: First audible agent reply timestamp: Endpointing premature? yes/no: Unexpected external action? yes/no/not enabled: Submitted audio duration / open session duration / actual billed units: Failure note: 5. Choose a configuration Your result table must name exact dataset, language, route, date and model. Keep vendor-published results separate from your own observations. Compare one change at a time. Review failures and raw sample counts. Winner for this defined workflow (if evidence supports one): Unresolved language / account / billing / carrier prerequisites: Release owner and human rollback plan: Burki setup reminders Use Standard pipeline for separate STT model selection. AssemblyAI is BYO; set speech recognition to BYO in Usage & billing and use the right account key. The latest assistant spoken reply is automatically shared with AssemblyAI to help recognize the next caller turn. GPT Live handles recognition itself and does not use these STT selectors. Provider selections and saved settings are not proof of a successful caller interaction. No tests or paid calls were run to create this worksheet.