GPT-LIVE 1 / GEMINI 3.8 LIVE BACKGROUND-WORK EVALUATION Reviewed October 9, 2026. Free worksheet, not free production voice usage. Capability comparison only. No matched model benchmark or provider calls run. Do not put secrets/customer data in this worksheet. Start with fictional read-only stubs; a stub checks app logic and is not a voice-model measurement. 1. EXACT CONFIGURATION Candidate A voice ID: gpt-live-1 Backend model/service, mode and prompt/version: Candidate B voice ID: gemini-3.8-live (base, not Extended Thinking) Tool declaration behavior and result scheduling: Current Google model/tools docs disagree on legacy defaults/interrupt enum spelling. Verify actual accepted SDK/API contract before configuring. Worked script is English only. Verify exact model language contracts; evaluate caller languages/accents in separate labeled sets. Modalities are not accuracy. Each: SDK/API version, transport, region, language, microphone/telephone path, prompt/policy version, tool implementation, network conditions, test date. Burki currently exposes GPT Live, not a verified Gemini Live selector. Current source uses client delegation; selector is not execution proof. 2. TASK STATE DESIGN TO REVIEW Task ID: Input revision and latest caller-confirmed reference: Tool call ID and input revision used: Pending/completed/cancelled/superseded state: Result provenance and external operation ID: Who can authorize a write: Duplicate prevention: How stale result is excluded from current spoken answer: Cancellation before start / while pending / after commit: Fallback when outcome is unknown: This is an engineering design worksheet, not an exposed Burki control. 3. FICTIONAL DELAYED READ Caller asks delivery request R204. Lookup starts with fixed local delay chosen before evaluation: ____ Caller corrects reference to R240 while pending. Old lookup returns: R204 awaiting collection. Expected: never attribute R204 status to R240; preserve correction; explain pending. Stub actual observed result (not model benchmark): Later authorized voice session/result evidence (if performed): 4. CASES AND PREDECLARED CRITERIA Case | Proposed pass | Observed result | Evidence | Reviewer Correction during read | no stale-reference claim: Cancel before write starts | write prevented: Cancel after external commit | state actual outcome/remedy, no false undo: Duplicate result/reconnect | no duplicate action: Tool timeout | outcome unconfirmed, agreed fallback: Playback interruption | stops audio as intended, task state accurate: Ambiguous 'stop' | clarify speaking vs lookup vs business action when needed: These are criteria, not measured scores. 5. COMPARABLE MEASUREMENT LOG Attempt | Candidate exact version/backend | Case | Delay/conditions Caller audio and transcript | provider acceptance/session ID | tool/task receipt Acknowledgment delay | time to correct useful answer | correction retained? Stale-result exposure? | duplicate action? | cancellation actual outcome Audio heard vs generated | failure/timeout | usage record | reviewer Keep all attempted cases, including failed connections and timeouts. If backend/config differs, disclose; don't call it voice-model-only benchmark. Report raw counts/sample size. Use repeated trials for tail latency; record the uncertainty method. No common current exact-pair benchmark verified here. 6. COST, AS LISTED OCTOBER 9 GPT-Live 1: USD 0.05 per session minute, billed per second; backend/tools separate. Illustrative six-minute voice-session component: USD 0.30, not an observed bill. https://developers.openai.com/api/docs/models/gpt-live-1 Gemini 3.8 Live paid per 1 million tokens: audio input USD 3, audio output USD 12; text input USD 0.75 / output USD 4.50. Provider also lists USD 0.005 per input audio minute and USD 0.018 per output audio minute. https://ai.google.dev/gemini-api/docs/pricing Session duration != audio output duration. Use recorded billable units. Add backend/tool/grounding/platform/telephone/transport costs as applicable. Candidate | actual billable units | rate/date | subtotal | other costs No savings ratio/winner inferred from unlike units. 7. BURKI INSPECTION AND DECISION Assistants -> draft -> Voice -> Model and provider choices if collapsed Conversation engine -> GPT Live option Record actual live model and separate reasoning configuration shown: Instructions/actions/readiness/credentials/Usage & billing review: If authorized testing occurs, record actual accepted session + caller-heard result + business receipt. Neither source nor dropdown proves acceptance. Selected configuration / evidence / unresolved limits / owner / review date: Primary contracts: https://developers.openai.com/api/docs/guides/live https://developers.openai.com/api/docs/guides/live-delegation https://ai.google.dev/gemini-api/docs/models/gemini-3.8-live https://ai.google.dev/gemini-api/docs/live-api/tools