Call Center Quality Assurance Scorecard: A Practical Template
Build a call center QA scorecard with clear criteria, a worked example and a fair way to handle missing evidence. Apply it to AI receptionist calls.
Table of Contents▼
A call center quality assurance scorecard is a repeatable way to review whether a conversation met your service standards. It turns “that sounded fine” into specific questions: Was the request understood? Were corrections retained? Did the caller leave with an accurate next step?
The same approach helps a small team reviewing an AI receptionist. You do not need dozens of weighted categories to start. You need a short rubric, evidence for each judgment and a way to distinguish a failed check from a check that cannot be judged.
Start with the type of call
An information enquiry, a callback request and a complaint require different outcomes. A caller asking for opening hours should not lose points because no sales lead was collected. A caller requesting a human should not be judged by whether the assistant kept the conversation going.
NiCE's guide to evaluation scorecards recommends matching the form to the interaction and keeping it understandable. Apply that principle by choosing one workflow first. The template below is designed for an enquiry that may require follow-up; adapt it before using it for other call types.
A call center QA scorecard template
Copy this table into your review document or spreadsheet. Add columns for the result, evidence and reviewer notes. This is a suggested manual worksheet, not a claim that Burki has a weighted scorecard builder.
| Check | Mark as met when | Evidence to review |
|---|---|---|
| Request understood | The response addresses the caller's final stated request | Relevant caller statement and response |
| Important correction retained | A corrected name, reference or contact detail replaces the earlier value | Correction and subsequent confirmation |
| Answer supported | The answer matches the approved business information available for that call | Answer and the relevant approved source |
| Limits explained accurately | The assistant states a relevant limitation without inventing access or authority | The explanation and available action evidence |
| Follow-up agreed | When follow-up is requested, the conversation clearly establishes the next step | The request and agreement |
| Closing accurate | The closing describes what happened and what remains open | Closing statement, compared with the rest of the record |
Use Met, Not met, Not enough evidence and Not applicable in this manual worksheet. The last category is useful when, for example, nobody corrected a detail or requested follow-up. It should not silently become a pass.
For each criterion, write down what would count as failure before reviewing calls. “Friendly and helpful” leaves too much room for interpretation. “Acknowledged the caller's stated concern before proposing the next step” gives reviewers something observable to check.
Worked example: a corrected service request
Imagine a fictional caller asking a property manager about a repair. They initially say unit 16, then correct it to unit 60. The assistant confirms unit 60, explains that it cannot check the repair's current status and agrees to pass on a callback request. The caller agrees. During closing, however, the assistant says, “Your repair is confirmed for tomorrow.” No appointment was arranged.
A reviewer could record:
- Correction retained: Met. The assistant confirmed the final unit number.
- Limits explained: Met. It accurately said that current repair status was unavailable.
- Follow-up agreed: Met. Both parties agreed on a callback request, without a promised time.
- Closing accurate: Not met. The closing invented a confirmed repair appointment.
That last failure matters even though several other checks passed. Flag unsupported commitments for correction rather than allowing an average score to conceal them. Review whether the callback request was actually saved separately; agreement during a conversation does not prove that another system received it.
Calculate a score without hiding uncertainty
If you need a percentage, start with a transparent manual calculation: met checks divided by met plus not-met checks. Report unknown and inapplicable counts alongside it. This is a suggested reporting convention, not a Burki calculation.
For example, three met checks, one not-met check, one unknown and one inapplicable check produce 75% across the four judged checks. Only four of the five applicable checks could be judged, so evidence coverage is 80%. Calling that simply “75% quality” would hide the incomplete evidence.
Do not compare teams or assistant versions until their call types, rubric and sampling approach are comparable. Keep any critical failure visible regardless of the percentage. A single invented commitment can require action even when the greeting, tone and contact details were all correct.
Apply success criteria in Burki
Burki's assistant settings include “What makes this conversation successful?” Use it to describe a concise outcome for the workflow. For a callback enquiry, a starting rubric might be:
The assistant acknowledged the final request, retained any corrected details and accurately explained the agreed next step without claiming an unconfirmed appointment.
This setting supplies a custom success criterion; listing several requirements does not create separate configurable weighted rows. Keep your manual worksheet if you need individual review dimensions or weights.
When saved analysis is available, the call's Call results section shows Quality checks. Outcomes appear as Met, Not met or Not enough evidence, with supporting references when provided. Burki's displayed outcomes do not include the worksheet's separate “Not applicable” category. Review conditional checks in context instead of interpreting every unknown as a failure.
Treat automated evaluation as a review aid. Check disputed judgments against the available transcript and, where needed, action records or audio. A text-based assessment cannot by itself establish whether the caller heard clear audio. The AI call summary guide explains the other saved outputs alongside these checks.
Calibrate before relying on the results
Have two reviewers independently assess a small, varied sample: ordinary enquiries, corrections, unclear requests, failed actions and calls that end early. Compare disagreements criterion by criterion. Decide whether the wording was vague, the evidence was incomplete or the judgment was wrong.
Revise the rubric and review the same examples again. Then test it on fresh calls. Keep the rubric version with your manual records so a changed definition does not look like a sudden performance improvement. Include difficult calls in later samples, not just long or successful conversations.
For an AI receptionist, start with one supported task and a few observable checks. Keep conversation quality separate from verified business completion. A useful scorecard tells the team what to fix next: the instructions, the available information, an action failure or the review criterion itself.
Ready to try Burki?
Create an assistant and check your available browser practice allowance.
Create your assistantTrial eligibility and available practice are shown in your workspace.