Zendesk AI agent QA means reading and scoring the conversations your Zendesk AI agent handles, against a written standard, so you know whether its answers were right and not just whether the customer went away. You need a scorecard built for the bot, a way to score every AI agent conversation, and a loop that turns each failure into a fix in Zendesk.
In short
- An automated resolution in Zendesk tells you the customer did not reach a human, not that the answer was correct.
- Start with the AI agent conversations where a wrong answer costs money: refunds, pricing, policy exceptions and account changes.
- Score the bot on accuracy, grounding, escalation and handoff, not on greetings and empathy alone.
- Every failure should end as a specific change: a help center article, a procedure, an escalation rule or a blocked topic.
- Re-score after each change, because a fix to one answer can break another.
This guide is the Zendesk version of our general guide on how to QA AI agents. If you are looking for the tool itself, the Zendesk QA tool page shows how Kaizo scores Zendesk conversations, including the ones Zendesk’s AI agents handle.
What does Zendesk show you about AI agent quality?
Zendesk shows you volume and outcomes: how many conversations the AI agent resolved, which ones it did not, and the transcripts behind both. It does not, on its own, tell you whether each answer was accurate or followed your policy. For that you need a review, either in Zendesk QA or in a separate QA tool.
In practice, a Zendesk admin has three places to look today:
| Where | What you get | What it does not tell you |
|---|---|---|
| AI agent insights and transcripts | Automated resolutions and unresolved conversations, with transcripts | Whether a resolved conversation was a correct answer |
| Automated resolution counts | How many requests the AI agent resolved without escalation to a human | Whether the customer came back, or gave up |
| Zendesk QA (paid add-on) | Reviews and autoscoring of conversations, including bots you mark as reviewable | Whether the answer followed your own policy, unless your scorecard checks for it |
Two details are worth knowing. Zendesk’s help center says the transcript view, which you open from the Insights tab of the AI agent’s settings, shows “the first 100 messages exchanged between an AI agent and a user” and that only conversations inactive for more than 72 hours appear among automated resolutions. The same article marks that view as legacy and says it will no longer be available after December 10, 2026, so check where your team will read transcripts after that date.
Why is an automated resolution not a quality score?
Zendesk measures AI agent usage in automated resolutions and bills “only for customer requests that were successfully resolved by an AI agent, without any escalation to a human agent.” That definition tells you the conversation ended without a human. It cannot tell you whether the answer was right, complete or allowed.
The gap shows up in three ways:
- The customer gave up. A wrong or generic answer, then silence. No escalation, so it counts as resolved.
- The answer was nearly right. The bot quotes the right policy but skips one condition, or the right product but the old price. Every system metric stays green.
- The bot promised something you do not offer. A refund outside the window, a fee waived, a delivery date nobody can meet. The customer is happy until the promise fails.
Support teams describe the same pattern in public. In one r/Zendesk thread, a team reviewing tickets “started and finished with the AI Agent (no human touch)” wrote that they saw “high value customers getting iffy or bad answers and then abandoning the chat”. That is a claim from one team, but it is the exact case an automated resolution count cannot show. We cover the wider list of these quiet errors in silent failures in AI agents.
How do you set up QA for Zendesk AI agents?
Set up QA for Zendesk AI agents in six steps: pick the risky topics first, write a scorecard for the bot, score every AI agent conversation, check the handoff to your human agents, trace each failure to a fix in Zendesk, and re-score after every change. The first two steps are a one-off job; the rest is a weekly routine.
Step 1: Decide which AI agent conversations matter most
Not every bot conversation carries the same risk. A wrong answer about opening hours is annoying. A wrong answer about a refund, a cancellation fee or a data request costs money or creates a compliance problem. List the topics where a wrong answer hurts, then make sure your review covers those first:
- Refunds, credits, compensation and fee waivers
- Pricing, plans and anything with a number in it
- Policy exceptions and eligibility rules
- Account changes, cancellations and personal data
- Complaints and threats to churn
Also flag conversations from high-value accounts and any conversation where the same customer comes back within a few days on the same topic.
Step 2: Write a scorecard for the AI agent
Your human scorecard is a starting point, not the answer. Greeting and spelling rarely fail for a bot. Accuracy, grounding in your own content and knowing when to stop are where it fails. Keep 6 to 8 criteria, each one something a reviewer can check against the transcript. The next section has a worked example, and our guide to the AI agent scorecard explains which human criteria to drop.
Step 3: Score every AI agent conversation, not a sample
A bot tends to repeat the same mistake across conversations that touch the same topic, because it answers from the same content each time. A 2% sample may miss it entirely, or find it once and treat it as noise. Scoring every AI agent conversation shows the pattern: one article, one procedure or one rule behind dozens of bad answers.
With Kaizo you connect Zendesk, pick a scorecard and switch on AutoPilot. It scores each conversation against your criteria when it resolves and attaches a comment that explains the score to the ticket. You can scope which queues, channels or teams it scores.
Step 4: Check the handoff to your human agents
When the AI agent escalates, two things can go wrong: it escalates too late, after the customer is already frustrated, or it hands over without the context, so the customer repeats everything. Score both on the bot’s side (did it escalate on the right trigger, did it pass a summary) and on the agent’s side (did the agent read it). Our guide on scoring the handoff covers both directions.
Step 5: Trace each failure to a fix in Zendesk
A QA score that does not change the bot is just a report. For every failed criterion, decide what caused it and where the fix lives:
| What failed | Likely cause | Where to fix it |
|---|---|---|
| Wrong or outdated fact | The source article is wrong, old or missing | Update or add the help center article the bot answers from |
| Right policy, missing condition | The article buries the condition, or two articles disagree | Rewrite the article so the condition is in the answer itself |
| Promised something not allowed | Nothing stops the bot on that topic | Hand that topic to a human, or stop the bot from answering it |
| Escalated too late | The conditions for handing over are too narrow | Widen them: topic, signs of frustration, the same question asked twice |
| Handoff without context | No summary passed to the agent | Change the handoff so the agent gets a summary of the conversation |
Group failures by cause before you fix anything. Ten bad answers that come from one article need one fix, not ten.
Step 6: Re-score after each change
Content changes have side effects. A rewritten refund article can fix refund answers and break the related cancellation answers. After each change, look at the scores for the affected topic over the next week and compare them with the week before. If you score every conversation, that comparison is already there.
What does a Zendesk AI agent scorecard look like?
A Zendesk AI agent scorecard checks whether the bot’s answer was correct, grounded in your own content, allowed by policy, and handed over at the right moment. It needs fewer soft-skill criteria than a human scorecard and more criteria about facts and escalation. Here is a worked example for a retail support team.
| Criterion | Question the reviewer answers | Type |
|---|---|---|
| Correct answer | Was every fact in the reply true for this customer’s case? | Pass / fail, auto fail on money topics |
| Grounded in our content | Can each claim be traced to a help center article or a procedure? | Pass / fail |
| Policy respected | Did the bot stay inside refund, fee and eligibility rules? | Pass / fail, auto fail |
| Understood the request | Did it answer the question the customer asked, not a similar one? | 1 to 3 |
| Escalated when it should | Did it hand over on the right trigger, before the customer repeated themselves? | Pass / fail |
| Handoff context | Did the human agent get a summary they could act on? | Pass / fail |
| Tone | Was the tone right for an upset customer? | 1 to 3 |
And here is the kind of comment a reviewer, or AutoPilot, should leave on the ticket:
Policy respected: fail. The customer asked for a refund on day 34. The AI agent confirmed the refund, but the returns policy allows 30 days. Source: help center article “Returns and refunds”, which states the 30-day window only in its last paragraph. Suggested fix: move the window into the first sentence of the article and add an escalation rule for refund requests after day 30.
A comment like this is what makes the score useful: it names the criterion, the evidence in the transcript, the source of the error and the fix. Before you trust any automated score, run it against conversations your team has already rated and see where it disagrees. Our guide on how to validate AI QA scoring walks through that test.
What results can you expect from full coverage?
The published results come from QA on support conversations in general. At UiPath, Kaizo automated 100% of QA with 200% ROI and an 8% lift in quality score. At SteelSeries, a 30-person support team that started out on Zendesk, AI summaries halved the time QA took without reducing what it caught.
Neither result is about AI agent conversations alone. The mechanism is the same, though: once every conversation is scored, a recurring failure stops being an anecdote and becomes a number you can fix and track.
See it on your own Zendesk. Kaizo scores every Zendesk conversation against your own criteria, including the ones Zendesk’s AI agents handle, and attaches the explanation to the ticket. Book a demo.
Should you use Zendesk QA or a separate QA tool?
Use Zendesk QA if you already pay for the add-on and its scorecards fit how you want to review the bot. Use a separate QA tool if you want one scorecard and one set of reports across human and AI agent conversations, built around your own criteria. Both can work; the scorecard and the follow-up matter more than the tool.
What Zendesk QA offers for bots, according to Zendesk’s help center:
- You mark each bot as reviewable or not. Reviewable bots are included in autoscoring and manual reviews.
- You can filter conversations by participant, bot, bot type (workflow or generative) and bot reply count.
- Its AutoQA categories are greeting, empathy, spelling and grammar, closing, solution offered, tone, readability and comprehension, and “agents and bots are evaluated separately”.
- It requires the Quality Assurance or Workforce Engagement Management add-on.
Questions to ask any tool, Zendesk QA included:
- Can it score the criteria that matter for a bot, such as accuracy against your own content and policy, not only the standard categories?
- Does every score come with an explanation tied to the transcript?
- Can you test it against your own past ratings before it scores live?
- Does it score every AI agent conversation, or a sample?
- Can the people who fix the bot see the failures grouped by cause?
If you are comparing the two directly, our page on the Zendesk QA alternative sets out the differences, and the Kaizo for Zendesk page shows how setup works.
When this is not for you
Formal AI agent QA is not worth it everywhere. If your Zendesk AI agent handles a few dozen conversations a week, read them yourself: a weekly hour with the transcripts and a short checklist will find most problems. If the bot only answers low-risk questions such as opening hours and order tracking links, a monthly spot check is enough.
It is also too early if you do not yet have a written standard for what a good answer looks like. Write that first, even as a one-page list, because no tool can score against a standard that does not exist. And if your help center content is out of date across the board, fix the content before you measure the bot: QA will mostly tell you what you already know.
Frequently asked questions
Can Zendesk QA review conversations handled by AI agents?
Yes. Zendesk QA can include bots you mark as reviewable in autoscoring and manual reviews, and you can filter conversations by bot and bot type. It requires the Quality Assurance or Workforce Engagement Management add-on.
Does Kaizo work with Zendesk AI agents?
Yes. Kaizo connects to Zendesk and AutoPilot scores every Zendesk conversation against your own criteria, including the ones Zendesk’s AI agents handle. Each score comes with a comment that explains it.
How many Zendesk AI agent conversations should you review?
Score all of them if you can, because a bot tends to repeat the same mistake across conversations about the same topic. If you review by hand, start with the risky topics such as refunds, pricing and policy exceptions, and read every one of those first.
Can you use the same scorecard for your human agents and your AI agent?
You can share the criteria that apply to both, such as correct answer and tone, which keeps the results comparable. The AI agent needs extra criteria for grounding in your content, policy limits and escalation, and needs less weight on greetings and spelling.
How long does it take to set up QA for a Zendesk AI agent?
Most of the work is choosing the risky topics and writing the first scorecard, which you do once. With Kaizo, connecting Zendesk and switching on AutoPilot takes days, not a connector project, and you can test the scoring against your past ratings before it goes live.
What should you do when the AI agent gives a wrong answer?
Find the source of the answer first: usually a help center article that is wrong, outdated or unclear. Fix the article or add an escalation rule for that topic, then check the scores for that topic over the next week to confirm the fix held.
Keep reading
- How to QA AI agents and chatbots, the general guide for any AI agent
- Your AI agent is closing tickets. Who is checking it?, on what to sample first
- QA inside the Zendesk ticket view, on why feedback belongs in the ticket
- Kaizo for Zendesk, the Zendesk QA tool for human and AI agent conversations