Use a Factual Answer Rubric for Voice Assistant Reviews
Review voice answers with a practical rubric for correctness, source support, completeness and uncertainty, while keeping serious factual failures visible.
Table of Contents▼
A voice assistant factual answer rubric makes review more consistent. Without one, a confident, friendly answer can receive a high score even when it omits a condition or invents a detail. Another reviewer may reject a cautious but correct response because it sounds less polished.
Separate factual quality from style. Both matter, but the business should know whether it is fixing an unsupported answer, an incomplete explanation or awkward delivery. A single “good call” score rarely provides that clarity.
Define the evidence available to the assistant
Before scoring, identify the approved documents, relevant tool results and caller information available at that turn. Review against that evidence, not against facts the assistant could not reasonably access.
A hypothetical service policy might allow a general explanation of inspection fees but require staff approval for a final quote. An answer that explains the process can be correct. A precise invented quote is not made acceptable by sounding helpful.
LiveKit's testing overview includes grounding and error handling as separate testing concerns. The rubric below is an operator-designed method for reviewing those concerns, not a built-in scoring feature promised by Burki.
Score four factual dimensions
Use a small scale with clear anchors. For each dimension, choose pass, needs review or fail, and retain the sentence that supports the decision.
| Dimension | Pass | Failure example |
|---|---|---|
| Correctness | Statements match approved evidence | Wrong opening time |
| Support | Claims come from an allowed source or confirmed result | Invented service guarantee |
| Completeness | Necessary conditions remain in the answer | Omitting a required eligibility rule |
| Uncertainty | Missing evidence is acknowledged accurately | Treating a failed lookup as a definite negative |
Keep pronunciation, concision and warmth on a separate style sheet. This prevents a pleasant voice from compensating mathematically for a material factual error.
Add an action-truthfulness gate
If an answer says something was done, require evidence that the action reached the claimed state. “Your request was saved” and “someone has completed the work” are different assertions.
A false action confirmation should be a distinct failure, even if the rest of the answer is accurate. The caller may act on that sentence. Do not average it away with several correct statements about less consequential details.
Where the outcome is genuinely uncertain, the assistant should say so and follow the supported recovery process. A cautious statement is not a failure merely because it avoids an attractive promise.
Calibrate reviewers on examples
Give two reviewers the same small set of conversations and ask them to score independently. Compare disagreements before using the rubric at scale. Often the problem is a vague business rule rather than the reviewers' judgement.
For the hypothetical fee policy, agree whether quoting a range is allowed and what qualifications must accompany it. Add that decision to the answer standard so future reviews use the same boundary.
Burki's voice testing guide can help organise these cases. Preserve clear failures and borderline cases; the latter are useful for making policy language more precise.
Review a balanced sample
Include ordinary questions, corrections, missing information and conflicting-source cases. Do not sample only successful calls or the conversations selected for a demonstration. Note the scope and size of the sample so a small review is not described as a universal accuracy rate.
When reporting results, show serious failures separately from minor omissions and uncertain reviews. If an automated evaluator helps triage the sample, retain human review for consequential or disputed cases rather than treating its score as ground truth.
Turn findings into controlled changes
Link each failure to a concrete repair: update a source, clarify a prompt, fix a tool result or narrow an unsupported claim. Rerun the affected cases and a small regression set after the change.
Burki's production-safe learning guide provides the review discipline. Start with the four factual dimensions and one action-truthfulness gate, then refine the examples until different reviewers can explain the same verdict.
Ready to try Burki?
Create an assistant and check your available browser practice allowance.
Start Free TrialTrial eligibility and available practice are shown in your workspace.