INWORLD TTS-2 / TTS-2 FLASH STEERING EVALUATION WORKSHEET Prepared October 11, 2026. Free worksheet; voice generation is metered. This is an unperformed protocol with fictional scripts, not benchmark results. Neither TTS-2 model ID is exposed in the verified native Burki selector. No voice generation, paid test or provider action is performed by this file. 1. DEFINE THE REQUIREMENT BEFORE SCORING Business message owner: ______________ Reviewer: ___________ Required outcome: __________________________________________ Is documented steering essential? __________________________ Is unsteered speech acceptable? ____________________________ Script approval / version: _________________________________ Caller languages / accent expectations / local names: _______ Permitted voice and permission evidence: ____________________ Direct provider access and budget owner: ___________________ Eligibility decision: - TTS-2 supports steering; Flash ignores steering instructions. - Do not give Flash a steering success score for pleasant neutral speech. - A supported control does not guarantee your desired listening result. 2. FIX THE COMPARISON CONTRACT Exact models: inworld-tts-2 / inworld-tts-2-flash API route and SDK version: _________________________________ Voice identifier for each model: ___________________________ Spoken language: ___________________________________________ Delivery mechanism: [ ] inline tags [ ] request instruction Instruction mechanism supported by chosen transport? ________ Encoding / sample rate / client region / playback: __________ Connection reused or newly opened: _________________________ Number of planned attempts, including failures: _____________ Fixed segmentation and context strategy: ___________________ Request version and actual metadata location: ______________ Budget ceiling and stop condition: _________________________ 3. SCRIPT A: NEUTRAL BASELINE, SAME SPOKEN WORDS The entrance is beside the library, not the loading gate. What would you like the team to clarify? Evaluate both models only after access, permission and budget are resolved. Listener should correctly repeat the entrance and understand the final question. 4. SCRIPT B: DIRECTED PASSAGE WITH EXPLICIT RESET [say slowly with clear articulation] The entrance is beside the library, not the loading gate. [reset] What would you like the team to clarify? Steering eligibility: TTS-2 only. Intended behavior, not measured result: - entrance instruction deliberately articulated; - question no longer depends on the earlier slow instruction; - all spoken words retained and intelligible. 5. SCRIPT C: SCOPE CONTROL, SAME WORDS WITHOUT RESET [say slowly with clear articulation] The entrance is beside the library, not the loading gate. What would you like the team to clarify? Compare B and C without changing model, voice, text or audio format. In a separate WebSocket exercise, put the final question in a later message within the SAME context. Keep reset/no-reset variants separately versioned. A new message is not automatically a new instruction scope. These are planned checks, not promised output or an installed SDK example. 6. OBSERVATION RECORD: ONE ROW PER ATTEMPT Blind label: ______ Exact model: ______ Script/version: ______ Request ID / retained evidence: _____________________________ Success/error and reason: ___________________________________ All words retained? ______ Entrance repeated correctly? ______ Unwanted delivery carryover? _______________________________ Question intelligible? ______ Voice consistent? _____________ Listener preference, kept separate from correctness: ________ First byte boundary / measured time: _______________________ First audible speech boundary / measured time: ______________ Last audible word boundary / measured time: _________________ Actual billable characters / charge: ________________________ Reviewer / language / date: _________________________________ 7. AGGREGATE WITHOUT INVENTING A WINNER Sample count including errors: _____________________________ Accepted messages / failed messages: ________________________ Missing words / misunderstood directions: ___________________ Timing boundary, unit, measurement location and percentile: __ Provider claims kept separate from your own measurements? ___ Total actual usage / accepted output: _______________________ Evaluation covers only English fixture, or other languages? _ Other accents/local names require new acceptance evidence: __ Vendor P90 first-byte figures are server-side and exclude network. They are not measured caller-heard latency. No shared independent latest-pair quality benchmark was established in the guide. Do not fill result fields from ads. 8. COST AND DEPLOYMENT DECISION Plan / currency / rate access date: _________________________ Current applicable character rate / units: __________________ Billable character-count receipt: __________________________ Retry, application, LLM, storage, platform and phone costs: __ Accepted-output cost and calculation: _______________________ Chosen workflow and evidence supporting it: _________________ [ ] Documented eligibility checked, runtime still untested [ ] Authorized evaluation completed with retained evidence [ ] Hold deployment because acceptance failed or is incomplete Primary references: https://docs.inworld.ai/tts/capabilities/steering https://docs.inworld.ai/tts/tts-models https://docs.inworld.ai/tts/synthesize-speech-websocket https://inworld.ai/pricing Original guide: https://burki.dev/blog/inworld-tts-2-vs-tts-2-flash