Skip to content

Template

QA Scorecard: Framework, Examples, Template and Scoring Maths

A QA scorecard template with real weightings and auto-fail rules, plus the sample-size maths every other guide leaves out.

· Updated · 16 min read

On this page

A QA scorecard turns quality into weighted, scorable criteria. See what to include, examples by channel, how to build and weight one, plus a free template.

A QA scorecard is the backbone of any quality program. Get it right and every conversation is graded fairly against what actually matters; get it wrong and QA becomes box-ticking. This guide covers what a QA scorecard is, real examples by channel, how to build one, how to weight and score it, and a free template to start from.

What is a QA scorecard?

A QA scorecard is a structured list of weighted criteria used to evaluate the quality of a customer support conversation. Each criterion defines part of what a good interaction looks like, and each conversation is scored against them to produce an overall quality score that drives coaching and reporting.

Think of the scorecard as the definition of quality made concrete. Instead of a vague sense that a conversation was “good,” the scorecard breaks quality into specific, measurable elements, was the issue resolved, was the tone right, was policy followed, and assigns each a weight. It is what makes QA consistent, coachable and defensible.

QA scorecard definition and structure

Why the QA scorecard matters

The scorecard is where most QA programs succeed or fail. A good one:

  • Creates consistency so every reviewer, and every AI evaluation, grades to the same standard.
  • Makes feedback specific so agents know exactly what to improve, not just their overall score.
  • Aligns quality with outcomes by scoring the behaviors that actually drive CSAT and resolution.
  • Stands up to scrutiny when an agent challenges a score or a regulator asks for evidence.

A bloated or vague scorecard does the opposite: it turns QA into subjective box-ticking that agents distrust and managers cannot act on.

How QA scorecards bring structure and consistency

What a good QA scorecard includes

Every effective scorecard has a few common ingredients:

  • Clear criteria written as objective, answerable questions.
  • Weights that reflect how much each criterion matters to your business.
  • A scoring scale, often yes/no, or a 1 to 5 range for nuanced criteria.
  • Auto-fail rules for compliance-critical behaviors.
  • Sections or categories that group related criteria for readable reporting.

QA scorecard categories and criteria

Most customer service QA scorecards organize criteria into four categories:

CategoryExample criteria
ResolutionWas the issue fully resolved? Was the information accurate? Was the right solution offered?
ComplianceWas the customer verified? Were required disclosures given? Was policy followed?
CommunicationWas the tone empathetic and professional? Was the response clear? Was the customer’s concern acknowledged?
EfficiencyWas the issue handled without unnecessary transfers, repeats or delays?

QA scorecard examples by channel

A scorecard should flex to the channel. Here are simplified examples.

Types of QA scorecard templates by channel

Chat / messaging QA scorecard

CriterionWeight
Issue resolved accurately30%
Response time and pacing appropriate15%
Clear, well-formatted, empathetic messaging25%
Correct process and verification followed20%
Proper close and next steps10%

Email / ticket QA scorecard

CriterionWeight
Complete, accurate resolution in fewest replies35%
Professional, on-brand tone and grammar25%
Compliance and policy followed25%
Clear structure and next steps15%

Voice / call QA scorecard

CriterionWeight
Issue resolved, correct information30%
Identity verification and compliance25%
Empathy, active listening, tone25%
Call control and efficiency20%

How to build a QA scorecard

  1. Start from your quality definition. List what a great conversation looks like for your team, then group those into categories.
  2. Write objective criteria. Phrase each as a clear question a reviewer or AI can answer consistently. Avoid vague terms like “good service.”
  3. Assign weights. Give more weight to what drives customer outcomes. Resolution and compliance usually carry the most.
  4. Set auto-fail rules. Decide which misses fail the whole conversation regardless of the rest.
  5. Pilot and calibrate. Score a batch with your team, compare results, and refine wording until scoring is consistent.
  6. Review regularly. Update the scorecard as products, policies and customer needs change.

Taking action with QA scorecard results

For a deeper build walkthrough, see our QA scorecard guide.

Free QA scorecard template

Start from a ready-made, weighted QA scorecard you can adapt to your team, across chat, email and voice. Grab our QA scorecard templates and customize the criteria and weights to match your quality definition.

Weighting and scoring a QA scorecard

Weighting is what turns a checklist into a meaningful score. A few principles:

  • Weight by impact. Criteria that affect resolution and compliance should outweigh cosmetic ones.
  • Keep the scale simple. Yes/no is easy to apply consistently; use a 1 to 5 range only where nuance genuinely matters.
  • Normalize to 100. An overall percentage is easy to trend and compare across teams.
  • Avoid over-weighting subjectivity. The more a criterion depends on reviewer opinion, the more calibration it needs.

The published cards disagree, and that is the point

There is no industry-standard weighting, and the vendors who publish theirs do not agree. Balto weights by section: greeting 10%, communication 25%, compliance 20%, resolution 25%, closing 20%. VereQuest publishes a 100-point behaviour card weighting problem-solving 20, empathy 15 and ownership 15, with greeting at 5 and using the customer’s name at 2. Five9 weights outcomes rather than behaviours: CSAT 30%, FCR 25%, QA score 20%, AHT 15%, adherence 10%.

Balto gives thirty points of a hundred to how the conversation opens and closes. Our channel cards above give that almost nothing. Neither is wrong. Weights are a business decision, not a best practice you can copy, so derive yours:

  • Start from what actually goes wrong. Pull your last 50 escalations, refunds and complaints and tag each with the behaviour that would have prevented it. If 30 of them trace back to wrong information given confidently, accuracy deserves more than 5 points.
  • Check the weights predict something. Score 100 conversations, then correlate the section scores against the CSAT or resolution outcome for those same conversations. If your tone score has no relationship to whether the customer came back satisfied, it should not carry 20 points. This is half a day in a spreadsheet and almost nobody does it.

The N/A trap that quietly breaks most scorecards

If a criterion does not apply to a conversation, remove its points from the denominator. Do not score it zero, and do not silently leave it in.

The arithmetic matters more than it sounds. An agent scores 68 points on a card where a 20-point criterion was genuinely not applicable. Dividing by the reduced total gives 68 out of 80, which is 85%. Dividing by the original 100 gives 68%. That is a 17-point swing produced entirely by a bookkeeping choice.

VereQuest handles this explicitly, noting that when N/A is used “the total % or total # of points is reduced by that amount”. Check what your tool does by default, because not all of them get it right. If you want the general form, Zendesk’s internal quality score is the sum of ratings divided by the maximum available score across categories. The important word is available, not possible.

Auto-fail criteria

Some misses are serious enough that the conversation should fail regardless of an otherwise strong score. Typical auto-fails include a skipped identity verification, an omitted legal disclosure, sharing incorrect regulated information, or a serious breach of customer trust. Defining auto-fails explicitly keeps your scorecard honest: a conversation that ticks every communication box but breaks compliance is not a good conversation, and the score should say so.

MaestroQA identifies three legitimate categories: compliance and security breaches such as exposing card data or skipping verification, knowingly giving wrong information, and abusive conduct. Three rules keep auto-fails survivable:

  1. Cap the list at five. If everything is critical, nothing is.
  2. Every auto-fail must be objectively verifiable. None should depend on a reviewer’s judgement of tone. Score tone; never auto-fail on it.
  3. Report auto-fails separately from the score. An agent with one auto-fail and nineteen strong conversations is a different problem from an agent averaging 60%. Blending them into one rolling average hides both, and one zero in a five-conversation month drops the average by 20 points on its own.

Documentation and compliance on a QA scorecard

What is a good QA score?

There is no universal pass mark, because it depends on how demanding your scorecard is. Most mature programs set a target band, often around 85 to 90% and up, and treat scores below it as a coaching trigger. Two things matter more than the headline number: the auto-fail criteria that override it, and the trend per agent and team over time, especially whether it correlates with CSAT. A single score means little; a rising trend that tracks happier customers means everything.

There is a harder reason not to fixate on the headline number, and it is arithmetic rather than philosophy.

How many conversations should you review?

Published guidance varies by more than tenfold, and none of it reconciles with the rest.

Call Centre Helper recommends 5 to 6 calls per agent, with Garry Gormley of FAB Solutions calling it the sweet spot between fairness and resource cost. HiveDesk recommends 5 to 10 per agent per month, rising to 10 to 15 for new hires in their first 90 days. OttoQA argues for around 65 per agent per month, though that is a vendor argument built on aggressive volume assumptions and it cites nothing. C2Perform names the right statistical parameters, 95% confidence and a 5% margin of error with Z = 1.96, then never prints the formula.

So here is the arithmetic. Treat each reviewed conversation as a pass or fail against your quality bar. The margin of error on an agent’s true pass rate at 95% confidence, assuming that rate is genuinely 90%:

Conversations reviewed per agentMargin of error
5plus or minus 26.3 points
10plus or minus 18.6 points
25plus or minus 11.8 points
50plus or minus 8.3 points
100plus or minus 5.9 points

These figures are computed rather than sourced: 1.96 multiplied by the square root of p(1-p)/n. Treating a scorecard total as a binomial pass rate is a simplification, and a deliberate one.

An agent reviewed five times who scores 85% has a true quality rate somewhere between roughly 59% and 100%. You cannot distinguish them from a colleague who scored 95%. To narrow that to plus or minus 5 points you would need about 139 reviews per agent per month.

What to do with that, honestly:

  • Do not rank agents against each other on small samples, or gate pay on a five-conversation average. The measurement cannot support it.
  • Do use small samples for coaching. Five reviews will not tell you an agent is in the bottom decile, but they will show you that this agent skips verification. Finding a specific fixable behaviour needs far less data than ranking people does.
  • Do aggregate upward. Five reviews per agent across a 40-person team is 200 data points at team level, which is plenty to spot a process or policy problem.
  • Do raise the sample where the stakes are high: new hires, regulated conversations, and anyone whose performance is formally in question.

Scoring every conversation automatically removes this problem, because there is no sample left to be unrepresentative. It introduces different risks, covered further down.

Calibration: keeping scores fair

Calibration is what makes a scorecard trustworthy. In a calibration session, several reviewers, and increasingly the AI, score the same conversations independently, then compare and reconcile differences. This surfaces ambiguous criteria, aligns interpretation, and builds the shared standard that makes agents accept their scores. Without calibration, a “score” is just one person’s opinion. With it, the scorecard becomes a fair, consistent measure, which is the entire point.

Most guidance stops there, without saying what a passing calibration looks like. Attach numbers to it. Every month, three to five reviewers independently score the same six conversations with no discussion, then measure two things.

Score spread. The gap between the highest and lowest total on the same conversation. A practical working threshold is 10 points on a 100-point card. Anything wider means the criterion definitions are ambiguous, not that a reviewer is wrong.

Agreement per criterion. Find which specific criteria your reviewers disagree on. It is almost never the binary ones. It is the judgement scales with undefined points. If you want a formal measure, Cohen’s kappa corrects agreement for chance, and the standard Landis and Koch bands put 0.41 to 0.60 at moderate, 0.61 to 0.80 at substantial and 0.81 to 1.00 at almost perfect (summarised here). Aim for 0.61 or better on any criterion carrying real weight. Below 0.41, the criterion is measuring the reviewer rather than the agent.

Fix the card, not the reviewer. When calibration finds disagreement, rewrite the criterion definition. Telling reviewers to try harder does not survive contact with the next ambiguous conversation.

What this looks like from the agent’s side

Scorecards are designed by managers and experienced by agents, and the gap between those two views is where quality programmes quietly fail. Support agents discuss their programmes publicly and in detail, and the same patterns recur.

Sampling feels like luck, because at these sample sizes it partly is. One agent describes performing well on the conversation types they are strongest at, then: “I followed my call flow to perfection on my new bookings and yet, they didn’t pull those.” (r/callcentres) That is not paranoia. It is a margin of error of about 26 points being experienced from the inside.

Unequal samples destroy the comparison. From a thread on automated QA: “This means my score might be based on 20 calls and my coworkers is based on 34.” (r/callcentres) If you rank agents at all, hold the sample size constant.

A single uncontrollable event can wreck a rolling average. “My dog barked ( he rarely does) and I got a zero on a call that otherwise would’ve been a perfect 100.” (r/callcentres) This is the auto-fail averaging problem in the wild.

Criteria that ignore the contact reason read as busywork. One agent, on being scored for rapport-building: “I work for an energy retailer, I don’t need to take someone on a journey when they ask to make a payment.” If your card applies the same criteria to a password reset and a bereavement notification, it will be resented and it will measure the wrong thing. Use channel-specific or contact-type-specific cards, or use N/A properly.

Reviewer error without a correction path poisons the whole programme. “3 of the 4 things that they said I didn’t ask, I absolutely did.” (r/callcentres)

The fix appears in that same thread, described by a QA analyst whose programme works. They send calls to each other “to make sure we’re being fair and accurate”, agents can dispute any score, and when the QA team is wrong they “own up to it and award the points back”. That is calibration plus an appeals process, and it is the difference between a scorecard people trust and one they merely endure.

So build the appeal in. Give agents the scored conversation, the criterion-level breakdown, a deadline to dispute, and a named person who decides. Track your overturn rate. If it is near zero, agents have probably stopped bothering. If it is above roughly 10%, your card is ambiguous.

QA scorecard vs QA rubric

The terms are often used interchangeably. A QA rubric usually refers to the detailed scoring guidance, what earns each score on each criterion, while the QA scorecard is the overall structured instrument, the criteria, weights and scoring together. In practice, a good scorecard includes rubric-level guidance so reviewers and AI interpret each criterion the same way.

Common QA scorecard mistakes

  • Too many criteria. An overloaded scorecard dilutes the signal that matters and slows scoring.
  • Vague wording. Subjective criteria produce inconsistent scores and agent pushback.
  • No weighting. Treating every criterion equally hides what actually matters.
  • No auto-fails. Letting a compliance breach pass because the rest scored well.
  • Set and forget. A scorecard that never updates drifts out of line with the business.

For a checklist of what not to do, see the do-not checklist for QA scorecards.

Manual vs automated scorecard scoring

A scorecard is only as useful as your ability to apply it. Scored by hand, even a great scorecard reaches under 5% of conversations. Scored by AI, as with Kaizo scorecards, the same scorecard is applied to 100%, consistently, freeing reviewers for calibration and coaching.

Manual scoringAutomated scoring
CoverageUnder 5% of conversations100% of conversations
ConsistencyVaries by reviewer and dayOne consistent standard
EffortHours of grading weeklyRedeployed to coaching

See how automated QA applies your scorecard to every conversation, and how AI coaching turns the results into agent development.

100%of conversations scored against your scorecard, versus the under-5% manual average

100%of QA automated at UiPath, with +8% quality score per quarter

75%less coaching-prep time at EverHelp across 16 domains

QA scorecards for AI-handled conversations

As AI agents start handling conversations, the same scorecard should grade them too, on one standard, so you can compare human and AI quality directly. The important caveat is neutrality: a vendor that sells its own AI agents cannot grade them against a scorecard impartially, while a neutral platform can. A well-built scorecard, applied by a neutral engine, becomes the shared quality bar for your whole operation, human and AI alike.

Versioning your scorecard

Your scorecard will change, and when it does you break your own trend line. Scores before and after are not comparable, and almost nobody warns you about this.

Three practices are enough. Version the card and stamp every evaluation with the version used. Never change weights mid-quarter. When you do change it, re-score 20 archived conversations under both the old and the new card so you know the size of the shift, then annotate the change on every chart that crosses it.

Review the card twice a year against your current escalation reasons. Criteria that made sense two products ago tend to survive far longer than they should.

The bottom line

The QA scorecard is the single most important artifact in a quality program. Keep it focused, weight it by impact, define auto-fails, calibrate it, and update it as your business changes. Then apply it to 100% of conversations with automation so the standard you worked hard to define actually reaches every interaction, not just a sample. That is how a scorecard stops being a checklist and becomes the engine of consistent quality.

Frequently asked questions

How do you compute a QA score?

Sum the points earned, then divide by the points that were actually available on that conversation, not the card’s headline total. If a 20-point criterion did not apply, the denominator is 80, not 100. An agent on 68 points scores 85% under the correct method and 68% under the wrong one.

How many conversations should I review per agent?

Five to ten per agent per month is the common recommendation and it is enough to find coachable behaviours. It is not enough to rank agents against each other or to gate pay, because at five reviews the margin of error is around 26 percentage points. If you need individual comparisons, either raise the sample substantially or score everything automatically.

How do I improve a low QA score?

Look at criterion-level data, not the total. A 72% caused by consistently missing verification is a ten-minute fix. The same 72% spread evenly across every criterion is a different problem entirely. If you cannot see which criteria are dragging the score down, your reporting is the thing to fix first.

How do I create a QA checklist?

Start from your last 50 escalations and refunds, tag each with the behaviour that would have prevented it, and let the frequency set your criteria. Cap the list at 20 items and write a one-sentence definition for every score point.

What is a QA scorecard?

A structured list of weighted criteria used to evaluate the quality of a support conversation and produce an overall quality score that drives coaching and reporting.

How do you build a QA scorecard?

Start from your definition of quality, write objective criteria grouped into categories, assign weights by impact, set auto-fail rules, then pilot and calibrate until scoring is consistent.

What is a good QA score?

Most teams target roughly 85 to 90% and up, define auto-fail criteria for compliance, and focus on the trend and its correlation with CSAT rather than a single number.

What is the difference between a QA scorecard and a QA rubric?

A rubric is the detailed scoring guidance for each criterion; the scorecard is the overall instrument of criteria, weights and scoring. Good scorecards include rubric-level guidance.

How many criteria should a QA scorecard have?

Fewer than you think. Focus on the criteria that predict customer outcomes. An overloaded scorecard dilutes signal and slows scoring.

Can a QA scorecard be scored automatically?

Yes. Automated QA applies your scorecard to 100% of conversations with AI, consistently, while humans handle calibration and coaching.

Do you need different scorecards for chat, email and voice?

Usually yes, at least different weights. The core categories stay the same, but criteria like response pacing or call control are channel-specific.

Sources

Margin-of-error figures on this page are computed rather than sourced, and are labelled as such where they appear. Practitioner quotes are from public r/callcentres threads, linked inline.

Apply your QA scorecard to 100% of conversations

Book a 30-minute demo and watch Kaizo score every conversation against your scorecard automatically and turn the results into coaching.

Book a demo

In Kaizo Scorecards A scorecard is where your standards stop being tribal knowledge. Build criteria in the words your business already uses, and every conversation gets measured against them. See Scorecards

On this page

See this on your own conversations

We will score a sample of your real tickets against your standards, so the example is yours.

Trusted by global support teams

  • Foot Locker
  • SteelSeries
  • Canva
  • GetYourGuide
  • Instacart