Skip to content

Guide

What Is an Auto-Fail in QA? Definition and Examples

What an auto-fail means, examples of failures that earn one, and how it differs from simply giving a criterion a heavy weight on your scorecard.

· 5 min read

Part of: QA Scorecard: Framework, Examples, Template and Scoring Maths

On this page

An auto-fail, or critical error, is a QA criterion that zeroes the whole evaluation when breached. It’s for failures nothing else can offset, such as a compliance breach or a data-handling error.

In short

  • A criterion qualifies only if one breach is severe and checkable.
  • A long auto-fail list is usually a weighting problem.
  • With automated scoring, the bigger risk is false positives.
  • Write auto-fails narrowly and trace each one to evidence.

Why some failures cannot be averaged

Most QA scoring is an averaging exercise. A scorecard breaks a conversation into criteria, each earns points, and the total describes how good the interaction was. That works because most quality signals trade off against each other, so a slow resolution can be balanced by an unusually clear explanation.

An auto-fail is the exception to that logic. It declares that one thing is not part of the average: if it happens, the evaluation is zero no matter what else was good. Teams also call it an automatic failure, a critical error, or a fatal error, and it usually sits in its own section at the top of the scorecard rather than as a weighted line item. The reason is that some failures are categorically different from a low score. An account change made without verifying the customer’s identity is not a mediocre interaction. It is one that should not have happened, and averaging it against a friendly tone misrepresents what took place.

When a criterion earns the power to zero a score

Because an auto-fail overrides everything else, it should be the hardest thing on a scorecard to qualify for. The short version of the test is that a single breach must be genuinely severe, objective enough that two reviewers always agree it happened, checkable against evidence in the transcript, and within the agent’s control. A criterion that fails any of those belongs in the weighted section with a high weight instead, where it still hurts the score and still shows up in coaching, but does not zero everything.

The full bar, with worked examples of which criteria clear it and which do not, is in our guide to auto-fail criteria in QA. The distinction to hold onto here is simple: the question is never “is this important?” Almost everything on a scorecard is important. The question is whether one instance is non-compensable.

Examples of auto-fail and non-auto-fail criteria

A few examples make the line concrete.

CriterionAuto-fail or weightedWhy
Identity not verified before an account changeAuto-failObjective, visible in the transcript, and severe on its own
Required disclosure or consent language omittedAuto-failThe words were either said or not, and the omission can invalidate the interaction
Sensitive data mishandledAuto-failBinary, evidence-backed, and not offset by anything else in the conversation
Rude or dismissive toneHeavily weightedSerious, but reviewers draw the line differently, so it fails the always-agree test
Missed upsell opportunityWeighted, lowA commercial preference, not a severe failure

Auto-fails under automated scoring

The definition does not change when a machine does the scoring, but the risk profile does. A grader can be wrong two ways on an auto-fail: it fires when it should not have, or it misses a real breach. Those are not symmetrical. A missed auto-fail is the same gap manual sampling already had. A false-positive auto-fail hands a zero to someone who did nothing wrong, and once the floor has seen the system get a fatal call wrong, every other score becomes arguable.

That asymmetry is why automated auto-fail criteria should be written narrowly and tuned for precision, and why every fatal score needs to point at the exact lines that triggered it. A score you can trace back to the evidence can be checked and, if wrong, corrected. A score you cannot trace is just an accusation with a number attached. This matters most when the platform scoring the conversation also sells the AI agents being graded, because it has an interest in the outcome. Kaizo does not sell its own AI agents, so an auto-fail it issues traces to the transcript and to nothing else. The full treatment of precision, appeals, and coaching after a zero is in the auto-fail criteria guide.

Frequently asked questions

What is an auto-fail in QA?

An auto-fail, also called an automatic failure or critical error, is a QA criterion that sets an entire evaluation to zero when it is breached, regardless of how the rest of the conversation went. It exists for failures that are not compensable, such as a compliance breach or a data-handling error, which good performance elsewhere cannot make up for.

What is the difference between an auto-fail and a weighted criterion?

A weighted criterion contributes points to an averaged score, so a failure lowers the total but does not destroy it. An auto-fail sits outside the average and zeroes the whole evaluation on a single breach. A criterion qualifies as an auto-fail only when one instance is severe, objective, evidence-checkable, and within the agent’s control. Otherwise it belongs in the weighted section.

What are examples of auto-fail criteria?

Common examples are failing to verify a customer’s identity before making an account change, omitting a required disclosure or consent statement, mishandling sensitive data, or making a prohibited commitment such as a refund the policy does not allow. Each is objective, visible in the transcript, and severe enough on its own to make the interaction unacceptable.

Can AI apply auto-fail criteria automatically?

Yes, and objective criteria such as whether identity was verified are exactly what automated scoring handles most reliably. The condition is that each criterion is verified against conversations reviewers have already scored, written narrowly to avoid false positives, and traceable to the evidence that triggered it, so a fatal score can be checked rather than taken on faith.

In Kaizo Scorecards A scorecard is where your standards stop being tribal knowledge. Build criteria in the words your business already uses, and every conversation gets measured against them. See Scorecards

On this page

See this on your own conversations

We will score a sample of your real tickets against your standards, so the example is yours.

Trusted by global support teams

  • Foot Locker
  • SteelSeries
  • Canva
  • GetYourGuide
  • Instacart