Skip to content

Template

How to Weight QA Scorecard Criteria Without Guessing

Equal weights quietly break most QA scorecards. Set criterion weights from evidence, test them against real scores before rollout, and keep them honest.

· 12 min read

Part of: QA Scorecard: Framework, Examples, Template and Scoring Maths

On this page

Set each criterion’s weight from three things: how much the failure hurts the customer, how much risk it carries for the business, and how much the agent controls it. If you can’t explain a weight, the score won’t survive its first dispute.

In short

  • Equal weighting claims every criterion matters equally. That is never true.
  • Agent control is the input everyone forgets, and agents resent it most.
  • Make critical criteria an auto-fail gate instead of a 60-point item.
  • Backtest new weights on scored conversations. If nobody changes band, it was cosmetic.

Why equal weighting quietly breaks a scorecard

Almost every QA scorecard starts with equal weights, because it feels fair and it takes no thought. It is neither. Equal weighting is a claim, and the claim is that forgetting a greeting costs the business exactly as much as processing the wrong refund.

Three things follow, and they arrive in this order.

The score stops discriminating. When ten criteria are worth ten points each, a serious failure and a cosmetic one both cost ten. Two agents end up on 90, one because they were curt and one because they gave the customer wrong information about their money. A number that cannot tell those two apart is not measuring quality.

Agents optimise for the cheap criteria. People respond to the scoreboard you give them. If tone and greeting are the easiest points to protect and the hardest to lose, that is where effort goes. This is not gaming, it is the scorecard working exactly as designed.

The score decouples from the business. Once quality scores stop tracking escalations, reopens and customer dissatisfaction, leadership quietly stops using them. The programme continues, the reviews continue, and nobody makes a decision with the output. That is the state most QA programmes are actually in.

A quick diagnostic. Pull your two highest-scoring agents from last month and your two lowest. If you cannot articulate a real difference in the customer experience they delivered, your weights are not carrying any information. Before touching weights, make sure the underlying criteria are sound, which is a separate job covered in what a QA rubric is and how to build one.

Remember

Equal weighting is not the absence of a decision. It is the decision that every criterion matters the same amount, made by default and never revisited.

The three weighting models, and when each one is right

There are only three models in practical use. The mistake is not picking the wrong one, it is picking one by accident and never revisiting it.

ModelHow it worksRight whenFails when
FlatEvery criterion carries the same pointsThe programme is new, you have fewer than about ten criteria, and you genuinely do not yet know what mattersYou have outcome data and are still ignoring it. Flat should be a starting position, not a resting one
Weighted by impactCriteria carry different point values reflecting customer and business impactThe programme is mature enough to have evidence, and reviewers can explain each weightWeights get set by whoever is loudest in the meeting, then never revisited
Deduction from a ceilingEvery conversation starts at 100 and loses points per failure, with severity setting the deductionMost conversations are fine and you are hunting exceptions, or the work is compliance-heavyThere is no floor, so one bad conversation goes deeply negative and distorts an agent’s average

How to set weights from evidence instead of opinion

Weight on three inputs, in this order. None of this is unique to support QA: the UK government’s manual on multi-criteria analysis works through the same problem of scoring options against criteria and then weighting those criteria, and its central warning is that the weights carry the value judgement whether or not anyone admits it.

Customer impact. How much worse is this customer’s day because the criterion failed? Wrong information about a refund outranks a missing sign-off, and it is not close.

Business and regulatory risk. What does the failure cost beyond this one conversation? Anything with a legal, financial or data-protection consequence sits at the top, and some of it should not be in the weighted pool at all.

Agent control. This is the one that gets skipped, and it is the one agents notice. If a criterion depends on a slow internal system, a policy the agent cannot bend, or a handoff from another team, weighting it heavily punishes people for your process. Either weight it low, or fix the process and stop scoring it.

The weights carry the value judgement, whether or not anyone admits it.

Paraphrasing the UK government’s multi-criteria analysis manual

The one hour of analysis that does most of the work

Take three months of conversations you have already scored. For each criterion, split them into the ones where it passed and the ones where it failed, then compare a downstream outcome across the two groups. Reopen rate is the best single choice because it is objective and it is in your helpdesk already. Escalation rate and dissatisfaction work too.

You are looking for the size of the gap. A criterion whose failure moves reopen rate by a large margin has earned weight. A criterion whose failure moves nothing has not, and it should either drop down the scale or come off the scorecard entirely.

Two honest caveats. This is correlation, not causation, so use it to rank criteria rather than to compute exact point values. And criteria that fail very rarely will not produce a readable signal, which is usually a hint that they belong in the auto-fail gate rather than the weighted pool.

Tip

Use reopen rate as your outcome variable. It is objective, it is already in your helpdesk, and unlike satisfaction it does not depend on anyone responding to a survey.

Auto-fail is a gate, not a heavy weight

Here is the single most common scorecard design error. A team decides that a customer data protection breach is unacceptable, so they make it worth 60 of the available 100 points.

That does not do what they think. An agent can breach data protection and still land at 40, and if the rest of the scorecard is generous they can finish the month with a respectable average. The behaviour the team called unacceptable is, arithmetically, survivable.

If a behaviour should void the interaction, it needs to be a binary gate that zeroes the score, sitting outside the weighted pool entirely. That is what an auto-fail is for. Three rules keep it usable:

  • Keep them few. Three to five. A scorecard with a dozen auto-fails is a scorecard where everything is critical and therefore nothing is.
  • Keep them unambiguous. An auto-fail must be answerable yes or no by two reviewers who have never spoken. “Was rude” is not that. “Disclosed account details without completing identity verification” is.
  • Require a second reviewer. Zeroing someone’s score is a serious act and it should cost the programme something to do it.

Once auto-fails are pulled out, the weighted pool gets easier to reason about, because everything left in it is a matter of degree rather than a matter of kind.

Attention

Making a critical criterion worth 60 of 100 points does not make it an auto-fail. The agent can fail it and still score 40, and with a generous scorecard still finish the month acceptably. If a behaviour should void the interaction, gate it. Do not price it.

The bimodal scorecard, and why agents stop believing it

There is a specific failure state worth naming, because it is common and it is fatal to trust. It happens when a scorecard combines a long list of auto-fails with heavy deductions on everything else. Agents describe the result the same way every time: every item is either an instant zero or costs sixteen points, so in practice the only outcomes are a perfect score or a disaster.

Look at your distribution. If scores cluster at the top and the bottom with very little in between, the scorecard has stopped measuring and started sorting. Two things follow. Targets like 95% become arithmetically unreachable for anyone who takes a hard conversation, so the people handling your worst tickets score worst. And because there is no middle, there is nothing to coach toward, only a pass or a punishment.

The fix is almost always the same: fewer auto-fails, and smaller deductions on the criteria that remain.

How to test new weights before you roll them out

Never publish a reweighted scorecard straight into production. Backtest it first, which takes an afternoon and prevents most of the damage.

Take a sample of conversations that have already been scored under the old weights. Rescore them under the new ones. Do not re-review them, you are testing arithmetic, not judgement. Then compare the distribution.

Three readings, and each one tells you something specific.

  • Almost nobody changes band. The reweight is cosmetic. You have spent political capital on a change that alters nothing, and you should either go further or leave it alone.
  • Almost everybody changes band. You have not adjusted weights, you have changed the standard. That may be correct, but it needs to be announced as a new standard rather than slipped in as a tuning change, or you will lose the room.
  • Some people change, and it is the wrong people. This is the useful case. If agents your team respects drop sharply, the new weights encode a belief about quality that you do not actually hold. Find the criterion causing it before you go further.

Then run one calibration session against the new weights with the whole reviewer group before anything is published. Weights change what reviewers argue about, and you want that argument to happen in a room rather than in an agent’s one to one. If you are working from a standard starting point, the QA scorecard templates are a reasonable base to reweight from rather than starting on a blank page.

Weighting when an AI does the scoring

Automated scoring does not change the principles above. It changes the stakes of getting them wrong, in two specific ways.

Errors stop being occasional and become systematic. Under manual review at the 3% sample most programmes run, a badly weighted criterion distorts a handful of conversations a month, and a good reviewer quietly compensates. With 100% coverage revealing trends that 3% sampling never could, that same bad weight is applied to every conversation, consistently, in the same direction. Consistency is the point of automation, and it is exactly what makes a weighting error compound.

A weight is only as defensible as its explanation. An agent told they lost 30 points will ask which 30, and why. If the system can only produce a total, the weights will not survive contact with the floor no matter how carefully you set them. Every deduction needs to name the criterion, quote the moment in the conversation that triggered it, and be disputable. That is the difference between a score and a verdict, and it is covered in more depth in how to validate AI QA scoring.

Two practical adjustments for automated scoring:

  • Keep deductions coarse. Multiples of five or ten. Fine-grained weights imply a precision the model does not have, and they invite arguments about the difference between a seven-point and an eight-point miss.
  • Weight the things a model reads reliably. Procedural and factual criteria are graded consistently. Judgement calls about tone carry more variance, so give them less of the total until you have measured how accurate AI QA scoring is on them.

Kaizo’s Auto QA applies the weights from your own scorecard per criterion and leaves the reasoning attached to the conversation, so a disputed deduction can be traced back to the specific exchange that caused it rather than argued from memory.

When to reweight, and when to leave it alone

Reweight on a schedule, not on a grievance. Quarterly, attached to a QA calibration cycle, is the cadence most teams can sustain. Between those points, collect the complaints rather than acting on each one.

Four signals that a weight has genuinely expired:

  • A criterion passes more than about 95% of the time for two quarters running. It is no longer discriminating between agents. Either the behaviour is now universal, in which case celebrate and retire it, or the criterion is too easy to satisfy.
  • A policy or product change moved what matters. New regulation, a new refund policy, a new channel.
  • The score stopped tracking outcomes. Re-run the analysis from earlier in this guide. If the gaps have flattened, the weights are stale.
  • Reviewers keep overriding the same criterion. When people quietly compensate for a weight, they are telling you it is wrong.

One rule that matters more than any of them: version the scorecard, and never retro-apply new weights to historical scores. An agent’s March score was produced under March’s rules. Recalculating it in July destroys the one thing a QA programme depends on, which is that the number meant something at the time it was given. Keep the old version intact, start the new version from a clean date, and say so out loud. This is also what keeps your internal quality score comparable over time instead of quietly drifting.

Frequently asked questions

How many criteria should a QA scorecard have?

Eight to fifteen for most support teams. Below eight the score is too blunt to guide coaching. Above fifteen reviewers cannot hold the whole rubric in their head, agreement between them drops, and reviews take long enough that coverage suffers. If you need more than fifteen, you probably need two scorecards for two different types of work rather than one long one.

Should every criterion have the same weight?

No. Equal weighting is a claim that every behaviour matters equally, which is never true, and it produces scores that cannot tell a serious failure apart from a cosmetic one. Flat weighting is a reasonable starting position for a brand new programme with no outcome data yet, but it should be replaced as soon as you can rank criteria by their effect on reopens, escalations or dissatisfaction.

What is the difference between a weight and an auto-fail?

A weight says how much a failure costs. An auto-fail says the conversation is void regardless of everything else. They are different mechanisms and should not be substituted for each other. Making a critical criterion worth 60 of 100 points is the most common design error, because an agent can fail it and still score 40. If a behaviour should void the interaction, gate it rather than weighting it.

How do I set weights when I have no data yet?

Rank criteria by customer impact and risk with your reviewers, group them into three bands of high, medium and low, and assign coarse point values to the bands rather than to individual criteria. Then commit to revisiting after one quarter, when you will have enough scored conversations to compare failure rates against reopens. Coarse and revisable beats precise and invented.

How often should QA scorecard weights change?

Quarterly at most, tied to a calibration cycle. Weights that move more often than that stop being a standard and become a moving target, and agents lose the ability to know what good looks like. Always version the scorecard when weights change, and never recalculate historical scores under new weights.

Put your own weights behind every scored conversation

Bring the scorecard you use today and a month of conversations you have already reviewed. We will show you what your current weights are actually rewarding, and what changes when every conversation is scored the same way.

Book a demo Explore Agentic Auto QA

In Kaizo Scorecards A scorecard is where your standards stop being tribal knowledge. Build criteria in the words your business already uses, and every conversation gets measured against them. See Scorecards

On this page

See this on your own conversations

We will score a sample of your real tickets against your standards, so the example is yours.

Trusted by global support teams

  • Foot Locker
  • SteelSeries
  • Canva
  • GetYourGuide
  • Instacart