Skip to content

Template

The AI Agent Scorecard: What It Needs That a Human's Does Not

Which criteria to add, which human ones to drop, and how to score an AI agent fairly without letting it mark its own homework.

· 6 min read

Part of: How to QA AI Agents and Chatbots

On this page

An AI agent fails differently from a person, so its scorecard drops the human-only criteria and keeps the ones about solving the problem. It adds checks for hallucination, unsafe escalation, staying in scope and admitting what it does not know.

In short

  • Tone warmth tells you little. Factual accuracy tells you almost everything.
  • Score every AI conversation, because automated failures repeat at scale.
  • The worst AI failures are silent and never escalate.
  • Don’t let the vendor that sold the agent grade it.

Why you cannot reuse the human scorecard

The instinct when you deploy an AI agent is to score it with the scorecard you already have. It is tempting because the goal looks the same: a good customer conversation. But the scorecard you built for people encodes assumptions about how people fail, and an AI agent breaks those assumptions in both directions.

A human agent will not usually invent a refund policy that does not exist, but they may be curt under pressure. An AI agent will almost never be curt, but it will state a fabricated policy with complete confidence. Scoring the AI agent on tone warmth tells you little, because tone is the thing it does most reliably. Scoring it on whether every factual claim it made was true tells you almost everything, because that is where it actually fails. The scorecard has to move its attention to where the risk now lives. Our full method for QA on AI agents and chatbots covers the workflow around this; the question here is narrower: what actually goes on the card.

Human criteria that stop making sense

Start by removing the criteria that measure something an AI agent does not have or does not vary on.

  • Tone and warmth as a proxy for empathy: an AI agent’s tone is set by its prompt and barely varies, so scoring it measures the configuration, not the interaction. Empathy still matters, but as whether the response was appropriate to the customer’s state, which is covered below.
  • Effort and initiative: criteria like going the extra mile assume a person choosing to do more. They do not map onto a system executing instructions.
  • Adherence to a script the human memorized: replaced by scope and policy adherence, which is a different and sharper check for a machine.
  • Personal development criteria: anything about the agent learning or improving belongs to the model owner, not to a per-conversation score.

Removing these is not lowering the bar. It is pointing the bar at the failures that are actually possible.

Criteria that only apply to an AI agent

These are the additions that make it an AI agent scorecard rather than a repurposed human one. Each one describes a failure a person rarely produces and a machine produces routinely.

CriterionWhat it checksWhy a human rarely triggers it
Factual accuracy of every claimNothing the agent stated was invented or wrongPeople hedge when unsure; models assert with equal confidence whether right or wrong
Policy fidelityThe agent did not invent, soften, or overstate a policyA human knows they do not know a policy; a model will generate a plausible one
Scope adherenceThe agent stayed within what it is allowed to do or promisePeople sense the edge of their authority; a model has to be told and can drift past it
Escalation behaviorIt handed off when it should have, and did not trap the customerA human recognizes being stuck; a model can loop or dead-end without noticing
Honesty about uncertaintyIt said it did not know rather than guessingGuessing confidently is a model default, not a human one
Safe handling of sensitive requestsIt refused or routed self-harm, fraud, or legal-risk cases correctlyJudgment a person applies instinctively has to be an explicit criterion for a machine

Criteria that stay the same

Some things matter regardless of who or what is answering, and these are the spine of the scorecard.

  • Was the issue resolved: the customer left with their problem solved, not merely without a human. This is the criterion a containment rate quietly skips.
  • Was the information correct and complete: shared with the human scorecard, but far heavier here.
  • Was the customer treated appropriately: reframed from warmth to appropriateness, whether the response matched the customer’s situation and emotional state.
  • Was the interaction efficient for the customer: not agent handle time, but whether the customer got there without unnecessary loops.

Writing a rubric an AI can score covers how to phrase each of these so it is checkable against the transcript rather than a matter of opinion.

How to score against it: coverage and traceability

Two things separate a working AI agent scorecard from a decorative one.

Score everything, not a sample

Human QA sampled a few conversations per agent because reading them was expensive. That logic breaks for AI agents, because their failures are systematic: a prompt weakness that produces one hallucination produces it hundreds of times, and a 2% sample will usually miss the pattern until it is already a problem. Scoring 100% of conversations is what makes the scorecard a monitoring instrument rather than a spot check. At UiPath, Kaizo automated 100% of QA with 200% ROI and an 8% lift in quality score.

Trace every score to the transcript

A factual-accuracy or policy-fidelity score is only useful if you can see the exact lines that failed it. Otherwise you cannot fix the prompt, and you cannot defend the score when the team that owns the agent pushes back. Every Kaizo score points at the evidence that produced it.

Grade the agent with something that did not build it

The uncomfortable structural point: if the platform scoring your AI agent is the same one that sold it, its scorecard is grading its own product, and the criteria most likely to be soft are exactly the ones that would make the product look bad. Kaizo does not sell its own AI agents. It scores the conversations they produce as a neutral party, which is the whole reason a criterion like factual accuracy or unsafe escalation can be scored honestly. The silent-failure taxonomy goes deeper on the failures that only a neutral grader tends to name.

Frequently asked questions

What is an AI agent scorecard?

It is the set of criteria used to evaluate conversations handled by an automated system rather than a person. It differs from a human QA scorecard because it drops criteria that only apply to people, keeps the ones about whether the customer’s issue was resolved and the information was correct, and adds criteria unique to automation such as factual accuracy, policy fidelity, scope adherence, and escalation behavior.

Can I use my existing QA scorecard for an AI agent?

Not without changing it. A human scorecard encodes assumptions about how people fail, and AI agents fail differently. Tone warmth barely varies for a model and tells you little, while fabricated facts and invented policy, which humans rarely produce, become the main risk. Reusing the human scorecard points your attention at the wrong failures.

What criteria are unique to scoring an AI agent?

Factual accuracy of every claim, policy fidelity (not inventing or overstating a policy), scope adherence, escalation behavior (handing off when stuck rather than trapping the customer), honesty about uncertainty, and safe handling of sensitive requests. Each describes a failure a human rarely produces and a model produces routinely.

Why should the AI agent scorecard not come from the vendor that built the agent?

Because that scorecard would be grading its own product, and the criteria most likely to be soft are the ones that would make the product look bad. A neutral grader that does not sell AI agents can score factual accuracy and unsafe escalation honestly, and trace each score to the evidence in the transcript.

Build an AI agent scorecard that catches what matters

Bring the conversations your AI agent handled last month and the scorecard you use today. We will show you which criteria stop making sense for a machine, which failures your current card cannot see, and what a scorecard built for automation catches instead. Because Kaizo does not sell its own AI agents, every score traces back to the exact lines in the transcript.

Book a demo Explore Agentic Auto QA

In Kaizo Kaizo Quality Assurance Kaizo scores 100% of your support conversations, human or AI, against your own quality criteria. Fully automated, or with human review where it matters. See Kaizo Quality Assurance

On this page

See this on your own conversations

We will score a sample of your real tickets against your standards, so the example is yours.

Trusted by global support teams

  • Foot Locker
  • SteelSeries
  • Canva
  • GetYourGuide
  • Instacart