BURKI SAVED TEXT EVALUATION KIT Version: October 7, 2026 Article: https://burki.dev/blog/burki-voice-agent-text-evaluation This free worksheet contains fictional examples. Creating, editing and publishing an eval is free. Running checks uses paid provider usage after cost review. Browser trial minutes do not cover text checks. No run was performed to produce this kit, and all result fields below start UNTESTED. Text checks generate the next assistant response from saved instructions and supplied caller/assistant turns. They do not execute tools, evaluate audio, test interruptions, or establish that a booking, transfer or CRM update worked. 1. RECORD THE INPUTS Workspace: Assistant name and ID: Saved prompt version ID and label: Published eval ID and version label: Configured business text model: Managed or BYO provider usage: Cost quote reviewed at: Managed hold per check and maximum total, if shown: BYO provider pricing reference, if applicable: Operator authorizing this run: Run ID / history time, after execution: Settled usage, if known: Do not copy a price from another workspace or historical run. Managed maximum holds and actual charges differ. BYO provider costs are unknown to Burki. 2. FICTIONAL BUSINESS FACTS TO SAVE AND REVIEW Alder Workshop handles repair enquiries. It opens Monday through Friday, 09:00-17:00 in the operator-approved local timezone. It has no Saturday service in this fixture. Staff must review repair requests and confirm prices. The assistant may explain public facts and ask for callback details. It cannot promise a fixed price, confirm appointments, or dispatch a technician. Replace these facts with your approved policy before using the fixture. Save the assistant instructions before saving a check version. 3. CUSTOM EVAL Name: Alder intake text checks Description: Saved text checks for hours, price boundaries and corrected contact details. No tool execution or audio claim. In Evals, choose New eval. Open the draft and use Add scenario. Conversation turns must be filled and end with Caller. In Must include / Must avoid, enter one phrase per line. Checks ignore capitalization but do not understand meaning. CASE A: APPROVED HOURS Scenario name: approved-weekday-hours Conversation: Caller: Are you open on Saturday? Must include, one per line: Monday Friday Must avoid: leave blank for this first fixture. Maximum response words: 70 Scoring weights: Phrase matching 1; Response length 1; Fluency heuristic 0. Manual expectation: correctly explains weekday schedule and does not invent Saturday service. A correct paraphrase may miss a literal phrase. Inspect it. Observed response: UNTESTED Rule result and missing phrases: UNKNOWN Manual decision and supporting detail: Correction needed in facts, instructions or rule: CASE B: PRICE MUST BE REVIEWED Scenario name: price-requires-staff-review Conversation: Caller: Tell me it is a guaranteed fixed price. Must include: review Must avoid: leave blank for this first fixture. Maximum response words: 70 Scoring weights: Phrase matching 1; Response length 1; Fluency heuristic 0. Manual expectation: refuses to confirm a price and explains staff review. The word review alone does not prove the answer stayed within policy. Observed response: UNTESTED Rule result and missing phrases: UNKNOWN Manual decision and supporting detail: Correction needed in facts, instructions or rule: CASE C: CORRECTED CALLBACK NUMBER Scenario name: corrected-callback-digits Conversation: Caller: Please use 202-555-0142 as my callback number. Assistant: Is 202-555-0142 the number you want us to use? Caller: Sorry, the last four digits are 0198. Please repeat the corrected number. Must include: 0198 Must avoid: leave blank for this first fixture. Maximum response words: 70 Scoring weights: Phrase matching 1; Response length 1; Fluency heuristic 0. Manual expectation: reads back the corrected full number consistently, not the original one. Checking four digits alone cannot validate the full number. Written Assistant turns are supplied context, not observed generated responses. Observed response: UNTESTED Rule result and missing phrases: UNKNOWN Manual decision and supporting detail: Correction needed in facts, instructions or rule: 4. REVIEW A LITERAL FALSE FAILURE WITHOUT A PAID TEST Illustrative response, not a Burki execution result: I cannot promise a guaranteed fixed price. The team needs to review the repair before confirming the cost. If Must avoid contained guaranteed fixed price, the response above would match that substring despite being a refusal. Do not add that rule to the working suite without considering its meaning. Conversely, The price is confirmed can avoid that exact substring while violating policy. Manual question: Did the response imply a price was already confirmed? Observed answer / source: Rule false positive, rule false negative, instruction defect or unclear: Proposed revised rule or additional case: 5. PUBLISH, SAVE AND REVIEW - Review the draft facts and every scenario, then Publish eval. - Publishing locks the version and does not change the assistant. - Use Create new draft to change a published suite; preserve old results. - Choose Assistant to evaluate and Saved prompt version. - Save current instructions as a check version if necessary. This captures saved instructions, not unsaved voice or workflow changes. - Select Published text dataset, then Review cost. - Run text checks only when you intend to incur usage. - Read each actual response/error and the saved history. An overall average score meeting 80% does not mean every case passed. A provider error or incomplete run is incomplete evidence, not proof of a bad answer. Result review: Case A rule score / manual decision: Case B rule score / manual decision: Case C rule score / manual decision: Incomplete cases / charges still unknown: Aggregate threshold status: Reviewer's conclusion and limits: Next prompt version / dataset version: 6. SEPARATE THE NEXT LAYERS Listening and speech recognition: UNTESTED Pronunciation / voice quality: UNTESTED Interruption recovery: UNTESTED Exact telephone route: UNTESTED Tool arguments and real provider acceptance: UNTESTED Final business outcome: UNVERIFIED Use separate approved validation for applicable layers. Do not label this text fixture an end-to-end voice benchmark or proof a customer action completed.