Skip to content

Best practice

QA Call Selection: Why "They Only Pick My Worst Calls" Matters

How to choose which calls QA reviews so the sample is fair, agents trust their scores, and coaching rests on representative evidence.

· 5 min read

Part of: What Is QA Sampling in Customer Service?

On this page

Call selection decides which conversations QA reviews, and so whether agents trust the scores. The common complaint that QA only picks an agent’s worst calls is usually an accurate read on selection bias.

In short

  • Reviewers drift toward short calls, flagged calls and calls near a complaint.
  • Scores from a skewed selection do not represent an agent’s real work.
  • Coaching built on those scores lands as unfair.
  • Fix it with truly random sampling, or score every conversation.

Why selection is where trust is won or lost

Every QA program that reviews a subset of conversations has to answer one question first: which ones? It is easy to treat that as logistics, but the answer quietly determines everything downstream. The conversations you select become the entire evidence base for an agent’s score, so if selection is skewed, the score is skewed, no matter how fair the scoring itself is.

This is also where agents form their opinion of the whole program. An agent does not see your scoring rubric or your calibration process. They see which of their calls got reviewed, and they notice the pattern. If the reviewed calls are consistently the hard ones, the program reads as a search for mistakes, and that perception, once formed, is very hard to reverse. A QA program’s credibility is decided at selection, before a single score is given. The wider trust question is covered in how to run a QA program agents trust.

Why manual selection is quietly biased

Reviewers rarely set out to pick unfairly. The bias creeps in through entirely reasonable habits, which is what makes it so persistent.

Selection habitWhy it seems reasonableThe bias it creates
Reviewing flagged or escalated callsThose calls seem most worth attentionOver-samples problems, so scores skew negative
Picking calls near a complaint or low CSATWanting to understand what went wrongReviews an agent at their worst moments
Choosing shorter callsThey are faster to get throughSystematically excludes complex work
Reviewing recent calls onlyThey are top of mindMisses patterns and rewards recency
Letting managers pickThey know their teamSelection reflects existing opinions of each agent

The agent complaint is usually correct

When an agent says QA only picks their worst calls, the instinct is to treat it as deflection. It is worth taking literally instead, because more often than not it is an accurate description of the selection habits above. If reviewers gravitate to flagged calls, complaint-adjacent calls, and escalations, then an agent’s reviewed sample really is weighted toward their difficult moments, and their score really does understate their typical work.

That has two costs. The obvious one is morale: coaching built on an unrepresentative sample feels like an ambush, and agents disengage from a process they experience as unfair. The less obvious one is that the data is genuinely wrong. You are making decisions about people based on a slice of their work that was selected precisely because it was atypical. The complaint is not just a feelings problem to manage. It is a measurement problem to fix.

The two honest fixes

There are only two ways to remove selection bias, and they sit at different points on the same line.

Make selection genuinely random and representative

If you must sample, sample properly: pull conversations at random across each agent’s full range of call types, lengths, and outcomes, rather than letting reviewers or managers choose. This is harder than it sounds, because true randomness has to be built into the process rather than left to good intentions, and even done well a small random sample carries the statistical limits covered in why a small sample cannot support agent-level decisions.

Remove selection entirely with full coverage

The cleaner fix is to stop selecting. If every conversation is scored, there is no sample to bias, and the argument about which calls got picked simply disappears. An agent’s score reflects all of their work, so the objection that QA cherry-picks the bad ones has no purchase. Scoring 100% of conversations is what makes that possible, and it changes the conversation on the floor from was this fair to what does the work actually show. At UiPath, Kaizo automated 100% of QA with 200% ROI and an 8% lift in quality score. Because every score traces back to the specific evidence, a coaching conversation starts from what happened rather than from a dispute about the sample.

Frequently asked questions

What is QA call selection?

It is how a QA program decides which conversations get reviewed. Because the selected conversations become the entire evidence base for an agent’s score, selection shapes both the results and whether agents believe those results are fair. It is one of the most consequential and most overlooked choices in a QA program.

Why do agents say QA only picks their worst calls?

Usually because it is true. Reviewers naturally gravitate to flagged calls, escalations, and calls near a complaint, which weights an agent’s reviewed sample toward their hardest moments. The result is a score that understates their typical work, so the complaint is an accurate read on selection bias rather than a defensive excuse.

How do you make QA call selection fair?

Two ways. Either make selection genuinely random and representative, pulling conversations across each agent’s full range of call types and outcomes rather than letting people choose, or remove selection entirely by scoring every conversation. Full coverage is the cleaner fix, because with no sample there is no selection to bias.

Does full coverage remove selection bias?

Yes, by removing selection. If every conversation is scored, there is no subset to skew, so an agent’s score reflects all of their work and the objection that QA cherry-picks bad calls no longer applies. It also shifts coaching from a dispute about which calls were chosen to a discussion of what the work actually shows.

Take the argument about call selection off the table

Tell us how you choose which conversations to review today. We will show you what scoring every conversation changes, how it removes the selection bias agents complain about, and how each score traces back to the evidence so coaching starts from what happened rather than a dispute about the sample.

Book a demo Explore Customer Service QA

In Kaizo QA automation Manual QA caps out at whatever your reviewers can get through. Kaizo scores every conversation against your own rubric, so coverage stops being a staffing question. See QA automation

On this page

See this on your own conversations

We will score a sample of your real tickets against your standards, so the example is yours.

Trusted by global support teams

  • Foot Locker
  • SteelSeries
  • Canva
  • GetYourGuide
  • Instacart