Back to Blog
Tutorials

Use a Factual Answer Rubric for Voice Assistant Reviews

Review voice answers with a practical rubric for correctness, source support, completeness and uncertainty, while keeping serious factual failures visible.

Burki
Article date:
4 min read

A voice assistant factual answer rubric makes review more consistent. Without one, a confident, friendly answer can receive a high score even when it omits a condition or invents a detail. Another reviewer may reject a cautious but correct response because it sounds less polished.

Separate factual quality from style. Both matter, but the business should know whether it is fixing an unsupported answer, an incomplete explanation or awkward delivery. A single “good call” score rarely provides that clarity.

Define the evidence available to the assistant

Before scoring, identify the approved documents, relevant tool results and caller information available at that turn. Review against that evidence, not against facts the assistant could not reasonably access.

A hypothetical service policy might allow a general explanation of inspection fees but require staff approval for a final quote. An answer that explains the process can be correct. A precise invented quote is not made acceptable by sounding helpful.

LiveKit's testing overview includes grounding and error handling as separate testing concerns. The rubric below is an operator-designed method for reviewing those concerns, not a built-in scoring feature promised by Burki.

Score four factual dimensions

Use a small scale with clear anchors. For each dimension, choose pass, needs review or fail, and retain the sentence that supports the decision.

DimensionPassFailure example
CorrectnessStatements match approved evidenceWrong opening time
SupportClaims come from an allowed source or confirmed resultInvented service guarantee
CompletenessNecessary conditions remain in the answerOmitting a required eligibility rule
UncertaintyMissing evidence is acknowledged accuratelyTreating a failed lookup as a definite negative

Keep pronunciation, concision and warmth on a separate style sheet. This prevents a pleasant voice from compensating mathematically for a material factual error.

Add an action-truthfulness gate

If an answer says something was done, require evidence that the action reached the claimed state. “Your request was saved” and “someone has completed the work” are different assertions.

A false action confirmation should be a distinct failure, even if the rest of the answer is accurate. The caller may act on that sentence. Do not average it away with several correct statements about less consequential details.

Where the outcome is genuinely uncertain, the assistant should say so and follow the supported recovery process. A cautious statement is not a failure merely because it avoids an attractive promise.

Calibrate reviewers on examples

Give two reviewers the same small set of conversations and ask them to score independently. Compare disagreements before using the rubric at scale. Often the problem is a vague business rule rather than the reviewers' judgement.

For the hypothetical fee policy, agree whether quoting a range is allowed and what qualifications must accompany it. Add that decision to the answer standard so future reviews use the same boundary.

Burki's voice testing guide can help organise these cases. Preserve clear failures and borderline cases; the latter are useful for making policy language more precise.

Review a balanced sample

Include ordinary questions, corrections, missing information and conflicting-source cases. Do not sample only successful calls or the conversations selected for a demonstration. Note the scope and size of the sample so a small review is not described as a universal accuracy rate.

When reporting results, show serious failures separately from minor omissions and uncertain reviews. If an automated evaluator helps triage the sample, retain human review for consequential or disputed cases rather than treating its score as ground truth.

Turn findings into controlled changes

Link each failure to a concrete repair: update a source, clarify a prompt, fix a tool result or narrow an unsupported claim. Rerun the affected cases and a small regression set after the change.

Burki's production-safe learning guide provides the review discipline. Start with the four factual dimensions and one action-truthfulness gate, then refine the examples until different reviewers can explain the same verdict.

Ready to try Burki?

Create an assistant and check your available browser practice allowance.

Start Free Trial

Trial eligibility and available practice are shown in your workspace.

Related Articles