A QA software pilot is a short, scoped test of a quality assurance tool on your own support conversations, with success criteria you write down before it starts. A good pilot runs two to four weeks, uses one or two queues, compares the tool’s scores with a reference set your reviewers agreed on, and ends with a yes or no you can defend.
In short
- A QA software pilot should run on your own conversations and your own scorecard, not on a vendor sandbox with demo data.
- Write the success criteria and the exit criteria before the pilot starts, so the result cannot be argued into a pass afterwards.
- Score a frozen reference set by hand first, then let every vendor score the identical set, or the numbers cannot be compared.
- Measure agreement per scorecard criterion, plus reviewer time saved and whether agents accept the feedback, not one blended accuracy figure.
- Name an owner, two reviewers and a decision date at the start; most pilots fail on missing people, not on the software.
What is a QA software pilot, and how is it different from a free trial?
A QA software pilot is a structured test with a fixed scope, a reference set and written success criteria, usually run with the vendor’s help. A free trial is open access to the product, often with demo data, that you explore on your own. A trial shows the interface; a pilot shows whether the tool scores your conversations the way your best reviewer would.
The difference matters because quality assurance software is judged on its output, not its screens. Three formats are common:
| Format | What you get | What it proves | Typical length |
|---|---|---|---|
| Free trial | Self-serve access, often a sandbox or a limited workspace | Whether the interface is usable | Days to weeks |
| Demo on your data | The vendor scores a sample of your real conversations and walks you through the result | Whether the scoring matches your standard on your tickets | One or two sessions |
| Pilot or proof of concept | A scoped run on live queues with your scorecard, reviewers and success criteria | Whether the tool works in your operation, with your people | Two to four weeks |
Trial lengths vary widely. SQM Group, which sells QA software itself, writes in its guide to questions for QA software vendors that trial periods “can range anywhere from 15 to 90 days”. Length matters less than scope: a 90-day trial with no reference set proves less than a two-week pilot with one.
At Kaizo we do not offer a sandbox trial. As our pricing page explains, we run the demo against a sample of your own conversations instead, because seeing your real tickets scored against your real standard tells you more than demo data would.
How long should a QA software pilot last?
Two to four weeks is enough for most support teams. One week goes to setup and the reference set, one to two weeks to scoring and comparison, and the last week to agent feedback and the decision. Longer pilots rarely add evidence; they usually mean nobody set a decision date.
Here is a four-week plan that works for chat, email and messaging queues.
| Week | Goal | What happens | Done when |
|---|---|---|---|
| 0 (before) | Agree the frame | Write scope, success criteria, exit criteria and the decision date; name the owner and reviewers | One page signed off by the budget holder |
| 1 | Connect and freeze | Connect the helpdesk, load your scorecard, pull and freeze a reference set of 60 to 100 conversations; two reviewers score it blind | The reference set and its agreed scores are locked |
| 2 | Score and compare | The tool scores the reference set and the live pilot queues; compare per criterion; diagnose every disagreement | An agreement table per criterion exists |
| 3 | Use it for real | Team leads use the scores and evidence in coaching; agents see results and can dispute them | Dispute rate and coaching notes are logged |
| 4 | Decide | Check each success criterion, collect the reviewer and agent verdicts, get the commercial terms | A written yes, no, or “extend with a named reason” |
The reference set method is the same one we describe step by step in our one-week protocol to validate AI QA scoring. This page does not repeat it; use that protocol for week 1 and week 2.
What should be in scope for a QA software pilot?
The scope of a QA software pilot should be one or two queues, the channels those queues use, your current scorecard, and a named group of agents and reviewers. Keep it small enough to finish in four weeks and representative enough that the result says something about the rest of the operation.
Choose deliberately:
- Queues: pick one high-volume, repetitive queue (billing, order status) and one harder queue (complaints, technical issues). A tool that only does well on easy tickets will look better than it is.
- Channels: include every channel the pilot queues handle. If your team works chat and email, test both.
- Scorecard: use the scorecard you run today, frozen for the pilot. If you know it needs work, test today’s version as a baseline. Our guide to what a QA rubric is covers the parts a scorecard needs.
- AI agents: if a bot or AI agent handles part of the queue, include its conversations. Many teams buy QA software partly to review what their AI agents say.
- Agents: a mix of tenured and new agents, so you see how the scores and the feedback land with both.
Leave out anything that does not change the decision: extra queues, a full rollout plan, every integration you might want next year.
What success criteria should a QA software pilot have?
Success criteria for a QA software pilot are the few measurable conditions that turn into a yes: agreement with your reviewers per criterion, reviewer time saved, coverage reached, evidence for each score, and agent acceptance. Write a target for each before the pilot, decided by the people who will live with the tool.
Use a table like this one and fill in your own targets. We do not give universal targets, because the right level depends on your scorecard and on how well your own reviewers agree with each other.
| Criterion | How to measure it | Who checks it |
|---|---|---|
| Agreement with your reviewers | Per scorecard criterion, against the locked reference set; precision and recall for pass or fail items | QA lead |
| Evidence per score | Share of scores that point to the exact part of the conversation behind them | Two reviewers, spot check |
| Coverage | Share of pilot conversations scored automatically, against your current sample | QA lead |
| Reviewer time | Hours per week on QA before and during the pilot | Reviewers, self-reported log |
| Agent acceptance | Dispute rate, overturn rate and what agents say in coaching | Team leads |
| Setup effort | Hours your team spent on connecting, configuring and fixing | Pilot owner |
| Fit with your stack | Helpdesk connection, user access, data residency and retention, exports | IT or security |
Then add exit criteria: conditions that end the pilot early with a no. For example, a criterion that matters for compliance where the tool misses failures your reviewers caught, or scores with no evidence you can check. Writing these down first protects you from a sunk cost decision in week four.
If the reviewers themselves disagree on a criterion, fix the wording before blaming the tool. The precision and recall behind an accuracy figure only mean something against a standard your own team agrees on.
Who should be involved in a QA software pilot?
A QA software pilot needs a pilot owner, two reviewers who score the reference set, one or two team leads who coach with the output, a small group of agents, someone from IT or security, and the budget holder. Each has a clear job, and the decision date is in everyone’s calendar from day one.
| Role | Job in the pilot | Time needed |
|---|---|---|
| Pilot owner (QA or support ops lead) | Writes the frame, runs the weekly check-in, owns the decision memo | A few hours a week |
| Two reviewers, plus a tie-breaker | Score the reference set blind, reconcile, diagnose disagreements | Most effort in weeks 1 and 2 |
| Team leads | Coach with the scores and evidence; report how agents react | Normal coaching time |
| Agents | Receive feedback, raise disputes, give a verdict | Normal work time |
| IT or security | Review the connection, access, data storage and retention | A short review early in week 1 |
| Budget holder | Signs off the frame in week 0 and the decision in week 4 | Two short meetings |
Agents are the group most often left out. A pilot that scores well but produces feedback agents reject will fail after rollout, so ask them.
How do you compare QA vendors fairly in a pilot?
You compare QA vendors fairly by giving every vendor the identical frozen reference set, the same scorecard version and the same success criteria, and by forbidding configuration changes after anyone has seen the comparison. Two accuracy figures measured on different samples cannot be compared, and a sample the vendor picked is not a measurement.
A few rules keep the comparison honest:
- You pick the conversations. Never let a vendor choose the sample it is measured on.
- Same rubric, same day. Load the same scorecard version into every tool before any scoring.
- No tuning after the reveal. Prompt or criterion changes after seeing the comparison fit the tool to your test set; if anything changes, rerun the whole set.
- Score per criterion. A blended number hides a criterion that does not work.
- Ask for the evidence. Every score should point at the lines that produced it. A score you cannot trace cannot be checked or disputed; see why QA score traceability matters.
- Compare total effort, not only licence cost. Count your reviewers’ and admins’ hours during the pilot.
When a named vendor is on your shortlist, our side-by-side pages on QA software alternatives list where each tool differs from Kaizo, for example the MaestroQA alternative and the Zendesk QA alternative pages. For a wider market view, start with our guide to call center QA software.
What should you ask a QA software vendor before and during the pilot?
Ask a QA software vendor how the pilot runs on your own data, who sets it up, how scores are explained, how agents dispute them, where data is stored, what it costs after the pilot, and what happens to your data if you do not buy. The answers show how the vendor will behave once you are a customer.
Questions worth asking, grouped by when you need the answer:
Before the pilot
- Will you score our real conversations, and can we choose the sample?
- Which helpdesks and channels connect natively, and how long does connecting take?
- Can we load our own scorecard, including weights and auto-fail criteria?
- Who on your side helps during the pilot, and how many hours do you expect from us?
During the pilot
- Can every score be traced to the exact part of the conversation behind it?
- How does an agent or reviewer dispute a score, and does the dispute change anything?
- Can the tool score our AI agent conversations with the same scorecard?
- What can we export, and in what format?
Before you sign
- How is the price calculated (per seat, per conversation, per scored item), and what changes it?
- What does onboarding include after the pilot?
- Where is our data stored, how long is it kept, and can we choose?
- What happens to the pilot data if we do not buy?
For the commercial side of the decision, our guide to the ROI of auto QA shows how to turn the pilot’s reviewer hours and coverage into a business case.
A worked example: one pilot on a support team of 40
A pilot frame for a 40-agent team fits on one page: the queues and channels in scope, a reference set of about 80 conversations, five success criteria with targets, two exit criteria and a decision date three weeks out. The team and figures below are illustrative, not a Kaizo customer.
A support team of 40 agents works chat and email in Zendesk. Today two QA reviewers score about 5 conversations per agent per month on a 10-criterion scorecard. The pilot frame they wrote in week 0:
- Scope: the billing queue (high volume) and the complaints queue (hard cases), chat and email, all 40 agents, the current scorecard frozen.
- Reference set: 80 conversations, stratified by queue and channel, scored blind by both reviewers and reconciled by the QA lead.
- Success criteria: agreement per criterion at or near the reviewers’ own agreement on the 7 objective criteria; no missed failure on the 2 compliance criteria in a separate set of known failures; every score linked to evidence; reviewer QA hours down; dispute overturn rate no higher than today.
- Exit criteria: any missed compliance failure that the reviewers flagged, or scores without evidence.
- Decision: written memo on day 20, signed by the head of support.
In week 2 the comparison showed two criteria where the tool and the reviewers disagreed most. The diagnosis found that one was a vague criterion (“showed ownership”) that the reviewers also scored differently. The team rewrote it, which improved their own scoring too. That is a common outcome: a good pilot improves the scorecard, whichever tool you buy.
What a pilot showed for teams that went ahead
When the pilot turns into a rollout, the gains come from scoring every conversation instead of a sample. At UiPath, Kaizo automated 100% of QA, with 200% ROI and an 8% lift in quality score. EverHelp, a BPO with more than 100 client projects, cut coaching prep by 75% after moving quality control off spreadsheets; the EverHelp customer story has the details.
See it on your own conversations. Kaizo scores 100% of your support conversations against your own scorecard and turns the findings into coaching. Book a demo and we will run it on a sample of your real tickets.
When a QA software pilot is not for you
A full pilot is more effort than it is worth when your team is small, your volume is low or you have no scorecard yet. A team of five handling a few hundred tickets a month can review a sample by hand in a spreadsheet, and a demo on your own data answers most of the questions a pilot would.
Be honest about where you are:
- No written scorecard yet: write one first. A tool cannot be tested against a standard that does not exist.
- Very low volume: manual review of a sample covers most needs. Our guide on when to stop doing QA in a spreadsheet lists the signs that you have outgrown it.
- No one free to run it: a pilot without an owner and two reviewers produces opinions, not evidence. Wait until you can staff it.
- A helpdesk Kaizo does not connect to: Kaizo connects natively to Zendesk and Salesforce today. If your conversations live elsewhere, check the integrations page before you plan a pilot with us.
When you are ready, the Kaizo QA platform page shows what the pilot would run on, and the pricing page explains how the plans differ.
Frequently asked questions
Do QA software vendors offer free trials?
Some do, often as a sandbox or a limited workspace. Others, including Kaizo, run a demo or pilot on a sample of your own conversations instead. A self-serve trial shows the interface; scoring your real tickets shows whether the tool matches your standard.
What is the difference between a pilot and a proof of concept?
The terms are often used for the same thing. When teams separate them, a proof of concept tests whether the tool can work at all on your data, usually on a fixed sample, and a pilot tests whether it works in daily operation with real reviewers, team leads and agents.
How many conversations do I need to test QA software?
A reference set of 60 to 100 conversations is a practical size for comparing scores per criterion. Add a separate set of known failures for rare compliance criteria, because a random sample may contain none of them.
Should agents know about the QA software pilot?
Yes. Tell agents what is being tested, that pilot scores have no consequences, and how to dispute a score. Their reaction to the feedback is one of the success criteria, and a pilot run in secret cannot measure it.
Can we pilot QA software on AI agent conversations?
Yes, if the AI agent’s conversations sit in a helpdesk the tool connects to. Kaizo scores conversations handled by Zendesk’s AI agents with the same scorecard as human agents, so a pilot can include both.
What happens after a successful QA software pilot?
Write the decision memo, agree the commercial terms, then roll out queue by queue with the scorecard you fixed during the pilot. Keep the reference set and rerun it every quarter, or whenever the scorecard or the scoring model changes.