Skip to content

Best practice

Scoring the Handoff: QA for Human-AI Transitions

What to score in both directions of the handoff, the failures that hide at the seam, and how to catch them when per-agent scores look fine.

· 5 min read

Part of: How to QA AI Agents and Chatbots

On this page

Handoff QA scores the moment a conversation passes between an AI agent and a human, in either direction. It is where blended support breaks most often, and per-agent scores miss it because the failure belongs to the transition.

In short

  • AI-to-human and human-to-AI handoffs fail in different ways.
  • AI-to-human fails on late escalation, lost context, or no usable reason.
  • Human-to-AI fails when an unresolved case goes back to automation.
  • Containment and handle time cannot see it. Read the whole conversation.

Why the handoff is its own failure point

When you QA a human agent and you QA an AI agent, you are scoring two halves of a conversation that increasingly is not two halves. In a blended model a customer might start with an AI agent, get escalated to a person, and be routed back to automation for a follow-up, all in one thread. Each participant can score well on their own stretch while the conversation as a whole is a bad experience, because the damage happened at the transition and belongs to neither of them.

That is the blind spot. A per-agent scorecard evaluates what each party did with the conversation while they held it. It has nothing to say about what happened at the moment they passed it on, which is exactly where blended support tends to break. The handoff needs to be scored as a thing in itself.

AI-to-human: what to score when the agent escalates

This is the more familiar direction, and it has three failure modes worth scoring explicitly.

Timing: did it escalate at the right moment

Too late is the common one. An agent that keeps trying after it is clearly stuck runs the customer in circles before handing over a frustrated person. Too early wastes the automation and the human’s time. The criterion is whether the escalation happened at the point the agent stopped making progress, which you can see in the transcript.

Context: did the human get what they need

The worst handoff makes the customer start over. Score whether the escalation carried the history, the diagnosis so far, and what the customer actually wants, so the human opens the case informed rather than cold. A handoff that drops context turns one conversation into two.

Reason: was the escalation intelligible

An escalation tagged only escalated to human tells the human nothing. Score whether the reason for the handoff was captured in a form the receiving person can act on. This is one of the criteria on an AI agent scorecard, but it only becomes visible when you read across the seam.

Human-to-AI: the direction nobody scores

The reverse handoff is newer and less examined, and it fails in ways that are easy to miss.

Human-to-AI handoffWhat good looks likeThe failure to score for
Returning a case to automation for follow-upThe AI has the resolution and the next step is genuinely routineThe human hands back an unresolved case to clear their queue
Delegating a sub-task mid-conversationThe AI can complete the specific task and hand back cleanlyThe AI cannot pick up the thread and the customer loses continuity
Post-resolution automated follow-upThe follow-up matches what was actually resolvedThe automation contradicts or forgets what the human just did

How to score a handoff in practice

The unit of evaluation is the conversation across the transition, not the segment on either side. That has three practical consequences.

Evaluate the whole thread

Score continuity across the seam: did context survive, did the customer avoid repeating themselves, did the second party pick up where the first left off. This is impossible if your QA looks at the AI portion and the human portion as separate records, which is how most tooling stores them.

Cover every handoff, because they are rare per pair but common overall

Any single agent pair may hand off infrequently, so a sample per agent misses handoffs almost by construction. Across the operation they are everywhere. Scoring 100% of conversations is what makes handoff failures visible as a pattern rather than as the occasional complaint. At UiPath, Kaizo automated 100% of QA with 200% ROI and an 8% lift in quality score.

Use a grader with no side in the transition

A handoff failure raises an awkward question: was it the AI’s fault for escalating badly, or the human’s for handing back a mess? A platform that sells the AI agent has an interest in the answer. Kaizo does not sell its own AI agents, so it scores the seam on the evidence, and every finding points at the exact turns where continuity broke, which is what lets you fix the routing rather than argue about blame. The full workflow is in how to QA AI agents and chatbots.

Frequently asked questions

What is handoff QA?

Handoff QA is evaluating the moment a conversation passes between an AI agent and a human, in either direction. It treats the transition as its own unit of evaluation, because a handoff can fail (lost context, late escalation, an unresolved case handed back) even when both the human and the AI score well on their own segments.

What should you score in an AI-to-human escalation?

Three things: timing (did it escalate at the point it stopped making progress, rather than too late or too early), context (did the human receive the history and diagnosis so the customer does not start over), and reason (was the escalation captured in a form the receiving person can act on).

What goes wrong in a human-to-AI handoff?

The main failures are a human handing an unresolved case back to automation to clear their queue, delegating to an AI that cannot pick up the thread and so breaks continuity, and post-resolution automated follow-ups that contradict or forget what the human just did. This direction is newer and rarely scored.

Why do per-agent QA scores miss handoff problems?

Because the failure lives between the records. A per-agent scorecard evaluates what each party did while holding the conversation, not what happened at the transition. Scoring the handoff requires reading the whole thread across the seam, which is why it needs the conversation stored and evaluated as one unit rather than two.

In Kaizo Kaizo Quality Assurance Kaizo scores 100% of your support conversations, human or AI, against your own quality criteria. Fully automated, or with human review where it matters. See Kaizo Quality Assurance

On this page

See this on your own conversations

We will score a sample of your real tickets against your standards, so the example is yours.

Trusted by global support teams

  • Foot Locker
  • SteelSeries
  • Canva
  • GetYourGuide
  • Instacart