Skip to content

Guide

What Is QA Sampling in Customer Service?

Why support teams sampled conversations for QA, what a small sample can and cannot tell you, and why full coverage is replacing it for agent decisions.

· 4 min read

QA sampling means reviewing a subset of conversations, often two or three tickets per agent per week, instead of all of them. It exists because manual review is slow, as a workaround for a capacity limit.

In short

  • A sample gives a rough team read, too thin for individual agents.
  • The hardest conversations are the least likely to be sampled.
  • Automated scoring removes the limit, so full coverage is replacing sampling.
  • Sampling still works for checking that reviewers or graders score consistently.

What sampling means in support QA

Sampling in customer service QA is simply the decision to review some conversations and not others. Because a reviewer can only read so many tickets or listen to so many calls in a day, a QA team picks a subset, scores those against the scorecard, and uses the result to represent the whole. A common shape is a fixed number per agent per week, chosen to fit the reviewers’ available time.

The word to notice is represent. The sampled conversations stand in for all the ones that were not reviewed, and the entire value of the exercise depends on how well they do that. When the sample represents the whole well, sampling is a reasonable economy. When it does not, every conclusion drawn from it is built on conversations that were never actually looked at.

One clarification, because the term travels across industries: sampling here means selecting customer conversations to review. It is unrelated to statistical sampling in manufacturing quality control, which is a different practice with a different purpose.

Why teams sample, and what it costs

Sampling is a response to a hard constraint. If your team handles thousands of conversations a week and you have a couple of reviewers, reading everything is impossible, so you read a slice. There is nothing wrong with the logic. The problem is what the slice can and cannot support.

A small sample gives you a rough sense of how a team is doing overall, and for that it is adequate. It struggles the moment you use it to judge an individual, because a few conversations is a tiny and often unlucky window on someone’s month. Two agents of equal quality can look very different on a small sample purely because of which conversations happened to be pulled. There is also a selection problem: the conversations that most need review, the messy escalations and the ones that went quietly wrong, are often the least likely to be sampled, because reviewers gravitate to conversations that are quick to score. The full statistical treatment of this is in why a small QA sample cannot support agent-level decisions.

Sampling vs full coverage

The constraint that made sampling necessary is going away. Automated scoring, such as Kaizo AutoPilot, can evaluate every conversation, which reframes the choice.

SamplingFull coverage
What gets reviewedA subset, sized to reviewer capacityEvery conversation
Good forA rough read on a teamAgent-level decisions and trend detection
Blind spotThe conversations never pulled, including the worstRequires trusting the automated scorer, so it must be verified
What the sample is forMeasuring qualityChecking the scorer, not measuring agents

Where sampling still belongs

Full coverage does not make sampling useless. It moves it to a better job. When every conversation is scored automatically, you no longer sample to measure agents, but you still pull a sample to check the measurement: to confirm that reviewers agree with each other and that an automated grader agrees with your reviewers. That is a legitimate, ongoing use of a sample, and it becomes more important under automation, not less, because a single grader now shapes every score. The related practices are QA calibration and validating AI QA scoring.

The move from sampling to coverage is also what lets QA stop being a source of argument on the floor. When an agent’s score rests on two conversations, it is easy to dismiss. When it rests on all of them, and every score traces back to the evidence, the conversation shifts from whether the sample was fair to what the conversations actually show. That is the case made in scoring 100% of conversations.

Frequently asked questions

What is QA sampling in customer service?

It is the practice of reviewing a subset of customer conversations instead of all of them, typically a few per agent per period. It exists because manual review is slow and there are always more conversations than reviewers can read. The sampled conversations are treated as representative of the whole, so the method is only as good as how representative the sample actually is.

Why is QA sampling a problem for judging individual agents?

Because a few conversations is a tiny window on an agent’s month, and which conversations were pulled matters enormously. Two agents of equal quality can post very different sampled scores by chance, and the hardest conversations, which most need review, are often the least likely to be sampled. A sample sized for reviewer capacity is too thin to support a decision about one person.

Is QA sampling the same as sampling in manufacturing?

No. In customer service, sampling means selecting conversations to review. It is unrelated to the statistical acceptance sampling used in manufacturing quality control, which inspects a subset of physical units against a defect tolerance. The two share a word but not a purpose.

Does full QA coverage replace sampling entirely?

It replaces sampling for measuring agents, because every conversation can be scored. It keeps sampling for a different job: checking the scorer, by confirming that reviewers agree with each other and that an automated grader agrees with your reviewers. That use of sampling becomes more important under automation, not less.

In Kaizo QA automation Manual QA caps out at whatever your reviewers can get through. Kaizo scores every conversation against your own rubric, so coverage stops being a staffing question. See QA automation

See this on your own conversations

We will score a sample of your real tickets against your standards, so the example is yours.

Trusted by global support teams

  • Foot Locker
  • SteelSeries
  • Canva
  • GetYourGuide
  • Instacart