Handoff QA scores the moment a conversation passes between an AI agent and a human, in either direction. It is where blended support breaks most often, and per-agent scores miss it because the failure belongs to the transition.
In short
- AI-to-human and human-to-AI handoffs fail in different ways.
- AI-to-human fails on late escalation, lost context, or no usable reason.
- Human-to-AI fails when an unresolved case goes back to automation.
- Containment and handle time cannot see it. Read the whole conversation.
Why the handoff is its own failure point
When you QA a human agent and you QA an AI agent, you are scoring two halves of a conversation that increasingly is not two halves. In a blended model a customer might start with an AI agent, get escalated to a person, and be routed back to automation for a follow-up, all in one thread. Each participant can score well on their own stretch while the conversation as a whole is a bad experience, because the damage happened at the transition and belongs to neither of them.
That is the blind spot. A per-agent scorecard evaluates what each party did with the conversation while they held it. It has nothing to say about what happened at the moment they passed it on, which is exactly where blended support tends to break. The handoff needs to be scored as a thing in itself.
AI-to-human: what to score when the agent escalates
This is the more familiar direction, and it has three failure modes worth scoring explicitly.
Timing: did it escalate at the right moment
Too late is the common one. An agent that keeps trying after it is clearly stuck runs the customer in circles before handing over a frustrated person. Too early wastes the automation and the human’s time. The criterion is whether the escalation happened at the point the agent stopped making progress, which you can see in the transcript.
Context: did the human get what they need
The worst handoff makes the customer start over. Score whether the escalation carried the history, the diagnosis so far, and what the customer actually wants, so the human opens the case informed rather than cold. A handoff that drops context turns one conversation into two.
Reason: was the escalation intelligible
An escalation tagged only escalated to human tells the human nothing. Score whether the reason for the handoff was captured in a form the receiving person can act on. This is one of the criteria on an AI agent scorecard, but it only becomes visible when you read across the seam.
Human-to-AI: the direction nobody scores
The reverse handoff is newer and less examined, and it fails in ways that are easy to miss.
| Human-to-AI handoff | What good looks like | The failure to score for |
|---|---|---|
| Returning a case to automation for follow-up | The AI has the resolution and the next step is genuinely routine | The human hands back an unresolved case to clear their queue |
| Delegating a sub-task mid-conversation | The AI can complete the specific task and hand back cleanly | The AI cannot pick up the thread and the customer loses continuity |
| Post-resolution automated follow-up | The follow-up matches what was actually resolved | The automation contradicts or forgets what the human just did |
How to score a handoff in practice
The unit of evaluation is the conversation across the transition, not the segment on either side. That has three practical consequences.
Evaluate the whole thread
Score continuity across the seam: did context survive, did the customer avoid repeating themselves, did the second party pick up where the first left off. This is impossible if your QA looks at the AI portion and the human portion as separate records, which is how most tooling stores them.
Cover every handoff, because they are rare per pair but common overall
Any single agent pair may hand off infrequently, so a sample per agent misses handoffs almost by construction. Across the operation they are everywhere. Scoring 100% of conversations is what makes handoff failures visible as a pattern rather than as the occasional complaint. At UiPath, Kaizo automated 100% of QA with 200% ROI and an 8% lift in quality score.
Use a grader with no side in the transition
A handoff failure raises an awkward question: was it the AI’s fault for escalating badly, or the human’s for handing back a mess? A platform that sells the AI agent has an interest in the answer. Kaizo does not sell its own AI agents, so it scores the seam on the evidence, and every finding points at the exact turns where continuity broke, which is what lets you fix the routing rather than argue about blame. The full workflow is in how to QA AI agents and chatbots.
Frequently asked questions
What is handoff QA?
Handoff QA is evaluating the moment a conversation passes between an AI agent and a human, in either direction. It treats the transition as its own unit of evaluation, because a handoff can fail (lost context, late escalation, an unresolved case handed back) even when both the human and the AI score well on their own segments.
What should you score in an AI-to-human escalation?
Three things: timing (did it escalate at the point it stopped making progress, rather than too late or too early), context (did the human receive the history and diagnosis so the customer does not start over), and reason (was the escalation captured in a form the receiving person can act on).
What goes wrong in a human-to-AI handoff?
The main failures are a human handing an unresolved case back to automation to clear their queue, delegating to an AI that cannot pick up the thread and so breaks continuity, and post-resolution automated follow-ups that contradict or forget what the human just did. This direction is newer and rarely scored.
Why do per-agent QA scores miss handoff problems?
Because the failure lives between the records. A per-agent scorecard evaluates what each party did while holding the conversation, not what happened at the transition. Scoring the handoff requires reading the whole thread across the seam, which is why it needs the conversation stored and evaluated as one unit rather than two.