ELEVEN V4 VS V4 TURBO COMPARISON WORKSHEET Burki free resource | Checked October 7, 2026 Original home: https://burki.dev/blog/eleven-v4-vs-v4-turbo No benchmark has been performed for this resource. Blank records are UNTESTED. Provider calls and synthesis can cost money. Arrange a bounded authorized test budget before running cases. Use fictional data and an agreed recording policy. 1. DEFINE THE DECISION BEFORE TESTING Workload (short assistant replies / produced narration / other): _____________ Language / listener population: _____________ Must-pass meaning/delivery cases: _____________ Maximum acceptable caller-to-audio delay and precise definition: _____________ Budget / approved owner: _____________ 2. FIX THE SHARED CONFIGURATION Model A: eleven_v4 | Model B: eleven_v4_turbo Provider account/plan/access review: _____________ Assistant draft: _____________ Standard pipeline: _____________ Same accessible voice ID for both models: _____________ Language: _____________ (start fixed; Automatic is a separate test) Recognition model/settings: _____________ Reasoning model/settings and approved business facts: _____________ Response text revision: _____________ Actual provider settings/transport/output format if visible: _____________ Network/connection conditions and location: _____________ Cold-start or reused connection: _____________ Voice/format incompatibility prevents identical setup? _____________ CURRENT UI LIMITS Voice -> Conversation engine: Standard pipeline ยท advanced provider choices -> Speech settings -> Advanced voice stack -> Voice provider: ElevenLabs -> exact Voice model. V4 Voice tuning shows delivery guidance and Language choices. It hides legacy sliders; do not invent Stability/Similarity adjustments in the current form. Provider v4 docs support Stability/Similarity, not Style/Speed/SSML. GPT Live separate provider choice controls Announcement voice only. 3. REUSABLE FICTIONAL SCRIPTS A. "We open at nine in the morning. I can explain the enquiry process." B. "Do you mean Wednesday morning or Thursday morning?" C. "I cannot confirm a booking from this enquiry. Staff need to review the request." D. "I can explain the available staff contact options." E. Longer approved answer containing an ordinary name, time and service limit. F. Plain response versus the same response with one permitted delivery tag. Keep tag trial separate from the plain baseline and identical across models. G. Authorized interruption case: caller corrects the question after reply begins. H. A language-specific case approved by a reviewer, after the fixed-language run. Keep actual task/action outcomes separate from synthesis quality. 4. MARK TIMESTAMP BOUNDARIES t0 = caller actually stops speaking (chosen observation method): _____________ t1 = synthesis work/request submitted (chosen receipt method): _____________ t2 = audible assistant output begins (chosen observation method): _____________ t3 = requested speech completes / is interrupted: _____________ Whole interaction delay = t2 - t0. Observed synthesis-plus-delivery interval = t2 - t1. Neither automatically equals a provider's model inference time. If timestamps are unavailable, write UNMEASURED instead of guessing from a recording feeling. Hold timestamp methods, output format and conditions constant across the pair. 5. CASE / ATTEMPT RECORD (copy for repeated attempts) Case ID / attempt / model / date: _____________ Exact text / voice ID / language: _____________ Cold or reused connection: _____________ t0 / t1 / t2 / t3 and units: _____________ Request accepted / rejected / no receipt: _____________ Actual playback / unavailable: _____________ Error / interruption / retry: _____________ Meaning understood by reviewer: YES / NO / UNTESTED Missed or altered words: _____________ Delivery suitability, with agreed scale: _____________ Model-blind reviewer ID and listening order: _____________ Business task/action result (separate): _____________ 6. SUMMARIZE WITHOUT CHERRY-PICKING Attempts per model/condition: _____________ Errors/retries/interruptions counted: _____________ Median measured interval and units: _____________ Percentile only if justified by sample count/method: _____________ Blind preference counts / ties / reviewer count: _____________ Unmeasured or non-comparable fields: _____________ Do not hide failures or call a small clip preference a production accuracy rate. 7. COST SNAPSHOT Provider pricing date / URL / plan / expiry: _____________ Unit rate and billable character definition: _____________ Characters in proposed workload: _____________ Synthesis-only estimate = characters / 1000 * unit rate. Separate platform, recognition, reasoning, telephony, taxes and plan allowances. October7 advertised rates through October12: v4 USD.022/1000chars, Turbo USD.011/1000chars; displayed ordinary rates USD.08 and USD.04. Recheck before use. These are a dated provider advertisement, not Burki quotes. 8. DECISION Choose / revise / defer: _____________ Evidence and important tradeoff: _____________ Must-pass failures still unresolved: _____________ Retest cases and review date: _____________ Do not infer a head-to-head winner from vendor Turbo-only latency figures. Voice selection: use Assistant voice. For Saved custom voice, expand Custom voice identifier and enter the actual permitted Custom voice ID.