Skip to content

Best practice

Peak Season Customer Service: How to Keep QA Honest

QA is the first thing cut when volume spikes, so visibility drops when risk peaks. How to resize sampling, pick what to relax, and score seasonal hires.

· Updated · 21 min read

Part of: How to Build a Customer Service QA Program: A 7 Step Playbook

On this page

Keep your QA programme running through peak, because reviewers get pulled onto the queue just as risk is highest. Decide the review budget, the sample and what you will relax before volume rises.

In short

  • A sample sized for normal volume doesn’t represent peak. Re-stratify it.
  • Write the relax and never-relax list before peak, not during it.
  • Give seasonal agents a shorter scorecard, not a softer one.
  • Book the post-peak review before peak starts.

Why QA is the first thing cut when volume spikes

Every peak season plan has a staffing section, a scheduling section and a self-service section. Almost none of them has a quality section. That is not an oversight in the writing, it reflects what actually happens on the floor.

Here is the mechanism. Your reviewers are usually senior agents or team leads, which means they are queue-capable. When volume doubles, the fastest lever a manager has is to put every queue-capable person back on the queue. QA review is the only support activity with no same-day customer consequence, so it is the cheapest thing to stop. Nobody announces the decision. The weekly review count just falls, then falls again, and by week three the programme has stopped without anyone choosing to stop it.

Three consequences follow, and they compound.

  • Measurement stops when the inputs change most. You are running the least familiar workforce you will run all year, on contact reasons that did not exist last quarter, for customers who buy from you once a year. That is the period you understand least well and the period you are now measuring least.
  • The record gets a hole in it. Peak is the season you will most want to study in January, when you plan next year. It will also be the season with the thinnest quality data, because you were not reviewing.
  • The programme restarts cold. Reviewers who have not scored anything for six weeks have drifted apart from each other, and the first scores of the new year are the least trustworthy ones you will publish.

Be honest about the tradeoff rather than pretending it away. An hour spent reviewing is an hour not spent answering, and you cannot run a full-fidelity manual QA programme and clear a doubled queue with the same people. Anyone who tells you quality and volume are compatible by default has not staffed a December. The useful question is what you give up, and whether you chose it in September or discovered it in December. This is the same constraint that makes understaffing expensive in ways that never show up on the staffing spreadsheet, and the same one that makes scaling a support team harder than adding headcount.

One practical note on timing. If you are a retailer, the window you are planning against is the November and December period the National Retail Federation tracks as the winter holiday season, which means the quality decisions in this guide need to be made in late summer, not in November when there is no room left to make them.

The peak decision timeline: what to decide, and when

Peak is the one support quality failure you can see coming. You already know roughly when volume rises, by how much, and which contact reasons arrive with it, because the same thing happened last year. That predictability is the opportunity, and most teams spend it: quality gets discussed in November, when every remaining option is a bad one, and reviewed in February, when nobody remembers what happened. Treat peak as a sequence of decisions with expiry dates instead. Each one is cheap and reversible at its own point on the timeline, and expensive or impossible after it.

WhenThe decision that belongs hereWhat it costs if you make it late
Eight weeks outWhether you run a quality programme through peak at all, and if you do, who is not going back on the queueNo budget, no hiring runway and no time to train a substitute reviewer. The decision gets made by whoever needs a body on Monday
Eight weeks outWhat you tell leadership will happen to the numbers, and what you are asking forThis is the only window in which a resourcing request can still get a yes, and the difference between a forecast dip and a discovered one
Two weeks outWhat is locked, what is frozen, and what the escalation trigger isChanges that land inside peak are the most common reason a programme collapses instead of shrinking
During peakNothing. You watch one signal and act only on the trigger you already agreedReopening the plan every week burns management attention and teaches the floor that the standard is negotiable
Two weeks afterWhat the season cost in quality terms, and the dated plan to repay itThe seasonal cohort has gone, the detail has faded, and the debt rolls into next year’s peak

Eight weeks out: run it, shrink it, or suspend it on purpose

There are three honest answers. Suspend deliberately, if your review capacity is one senior agent who is also your fastest closer. Shrink deliberately, keeping a much smaller programme against a much smaller scorecard aimed at the parts of the operation you understand least, which is the right answer for most teams. Or change the constraint, so that reviewing stops drawing on the same headcount that answers tickets, which is a procurement decision with a lead time. If you suspend, name the blackout window with a start and an end date, keep a minimal safety net on the criteria that would void an interaction, and accept that the season will leave a hole in your records.

Whatever you choose, translate it into names. “We will keep QA running” is a hope. “Priya and Marcus do not go back on the queue between 17 November and 5 January, and here is who covers their tickets” is a plan, and it is specific enough to defend in a staffing meeting.

Eight weeks out: brief leadership before the numbers move

A dip you predicted in September is evidence you understand your operation. The same dip surfacing in a January board pack is evidence you did not. Write one page and send it before anything moves:

  • Which numbers will move, in which direction and roughly how far. Ranges are fine.
  • Why, in one sentence a non-support executive can repeat. Usually: reviewers are queue-capable, so review capacity and answer capacity come out of the same headcount.
  • When it comes back, with a date. An open-ended decline reads as loss of control.
  • The version warning. If the scorecard changes for peak, scores produced under it are not comparable to the annual series. Say so now, because saying it in February sounds like an excuse.
  • The one thing you are asking for. A contract reviewer, two weeks of extra ramp, a tooling decision. One thing, not five. If it involves spend, frame it as avoided cost and rework, as in the ROI of automated QA, and see reporting support quality to executives for the wider translation problem.

Two weeks out: lock some things, freeze the rest

The two weeks before peak are for stopping things, not starting them. Lock, in writing and with dates: the peak scorecard version and its expiry date, the review commitment with names and backfill, the escalation trigger, and who owns quality if the QA lead goes back on the queue. Freeze until the season ends: new tooling (an implementation starting two weeks out will not finish, and the sequencing in a QA tool rollout plan will usually confirm it), new contact-reason taxonomies, new criteria and definitions, and anything that needs training for the whole floor. Two things get through the freeze because they expire: onboarding the seasonal cohort, and one calibration session on the rubric reviewers will actually use.

During peak: watch one signal, act on one trigger

The quality score is the wrong signal during peak. It lags by a week or more, it is produced by fewer reviewers than usual, and under a modified rubric it is not comparable to October. Pick something cheap, fast and already instrumented, usually reopen rate split by tenure cohort, or escalation rate by cohort if reopens are noisy. Then define the trigger before you need it: a measurable condition plus a pre-agreed response, such as “if reopen rate on agents under 90 days exceeds double the tenured rate for two consecutive weeks, one reviewer comes off the queue for three days and reviews only that cohort”. And do not reopen the relax list mid-peak. You wrote it calmly, with the whole picture in front of you.

What peak does to your sample, and why the old sample size lies

Most QA programmes review a small random percentage of conversations. That works when the population is stable. Peak is the one time of year it is not.

Four things change at once: who is answering, what customers are asking, who the customers are, and when the work happens. A random draw across the whole month will be dominated by your highest-volume contact reason handled by your highest-volume tenured agents, because that is what a proportional sample does. The sample will look healthy and will tell you nothing about any of the four things that changed.

The fix is stratification, not more reviews. Fix a review budget you can actually staff, then allocate it deliberately instead of drawing it at random.

A worked example

Say a normal month is 15,000 conversations and you review 3% of them, so 450 reviews. December is 40,000. Holding 3% means 1,200 reviews, which you cannot staff, and which is why the programme collapses instead of shrinking.

So go the other way. Decide the number of reviews you can genuinely complete, say 300, and spend them:

  • 120 on agents under 90 days. They may be 20% of headcount and less than 20% of volume, but they carry most of the variance.
  • 90 on peak-specific contact reasons. Delivery exceptions, gift orders, returns, promotion disputes. Reasons that barely existed in October.
  • 60 on business as usual, so you can still tell whether your baseline moved.
  • 30 on reopens and escalations, which are the cheapest early warning you have.

Three hundred well-chosen reviews beat 700 random ones, and 300 is a number a single reviewer can actually deliver alongside queue time. The point of QA sampling at peak is not statistical purity, it is making sure the parts of the operation you know least about are the parts you look at most.

What changes at peakWhy a normal-season sample misses itWhat to do instead
Who is answeringSeasonal and cross-trained agents are a small share of total volume, so a proportional sample barely touches themSet a minimum number of reviews per week for anyone under 90 days, independent of their share of volume
What customers ask aboutPeak-specific reasons are new, and a random draw over-samples the familiar top reason you already understandStratify by contact reason and reserve a fixed share of the budget for reasons that did not exist last quarter
Who the customers areOnce-a-year buyers do not know your policies or your product, and they escalate differently, but they are invisible in an agent-first sampleAdd a stratum for first-contact and first-purchase customers
When the work happensExtended hours and weekend cover fall outside the shifts reviewers work, so late and weekend conversations are systematically under-reviewedSample by time block as well as by agent, and check coverage by hour in week one rather than in January
How long a conversation isBacklog makes threads longer and multi-agent, so one score lands on whoever happened to touch it lastScore the handoff explicitly, or pull multi-agent threads out of individual scoring and review them as a process problem

Which criteria to relax, and which you never relax

Under a doubled queue, something gives. If you do not choose what, your agents will choose for you, and they will choose based on what gets measured and what gets shouted about in stand-up, which is almost always handle time.

There is one test for whether a criterion can be relaxed. Does failing it cost the customer something they will still notice after the conversation ends, or cost the business something it cannot undo? If neither, it is a candidate.

CriterionPeak statusWhy
Greeting, sign-off, brand voice, formattingRelaxThese are consistency criteria. A correct answer in four minutes beats a warm one in four hours, and peak customers say so
Proactive personalisation, upsell prompts, satisfaction-probing questionsSuspendDiscretionary value-add that assumes spare cognitive room. At peak there is none, so scoring it only manufactures failures
Documentation and taggingRelax the depth, hold the minimumFull notes are expensive. A ticket with no contact reason recorded is a ticket you cannot analyse in January, which is exactly when you will want to
Accuracy of the answer givenNever relaxA wrong answer at peak creates a second contact at peak, so it costs you twice at your most expensive moment
Identity verification, data handling, consent, regulated disclosuresNever relaxThese are gates, not weighted criteria. A gate that bends under pressure was never a gate
Resolution, ownership, correct escalation and handoffNever relaxReopens and transfers are the mechanism by which December backlog becomes February backlog

Put the peak scorecard in writing before peak starts

The difference between a plan and a panic is that the plan is written down and dated. A verbal understanding that the team is “being pragmatic about tone right now” is not a standard, it is deniability, and agents can tell the difference immediately.

Produce one page, before peak, and publish it to the whole team including the seasonal cohort on their first day. It needs five things:

  1. A start date and an end date. The peak scorecard expires. Say when, or it becomes the permanent scorecard by accident.
  2. The list of relaxed and suspended criteria, named individually, with one sentence of reasoning each. Reasoning is what stops it reading as management giving up.
  3. The gates, restated, so nobody reads “we are relaxing the scorecard” as “we are relaxing everything”. Anything that should void an interaction belongs in your auto-fail list, not in the weighted pool where it can be traded off against speed.
  4. The review budget and how it is allocated, so anyone can see why a seasonal agent is being reviewed four times a week and a tenured agent once a fortnight.
  5. An explicit note that peak scores are a separate version. Scores produced under peak rules do not average into the annual number as though they were the same measurement.

That last point is the one teams skip and regret. If your QA scorecard changes and the score keeps the same name, you have quietly broken the only thing a quality number is good for, which is comparison over time. Version it, label the period, and keep the two comparable to themselves rather than to each other.

How to score seasonal and temporary agents fairly

A rubric built for tenured agents mostly measures tenure. A three-week hire will lose points on product depth, tool fluency, policy edge cases and judgement calls, all of which are functions of time on the job rather than of effort or care. Score them against it and you get a number that predicts nothing, demoralises a workforce you cannot afford to lose, and drags the team average down in a way that misleads you about your tenured agents too.

Three adjustments fix most of it.

1. Cut the scorecard down, do not soften it

Reduce to the criteria a competent adult can meet in week one: correct answer, correct process followed, correct escalation, respectful tone, verification completed. Five to seven items. This is closer to a QA checklist than to a nuanced rubric, and that is appropriate. A short scorecard applied strictly is fairer and more useful than a long one applied with silent allowances.

2. Report the cohort separately, always

Never mix a seasonal cohort into the team average. Publish two numbers. Mixing them tells you the team got worse when what actually happened is that the team got bigger and newer, and it hides whether your tenured agents held up under load.

3. Compare against a ramp, not a target

Define what acceptable looks like in week one and what it looks like in week three, then compare an agent to their own cohort at the same tenure. A seasonal agent at 72% in week one who reaches 84% by week three is a success. The same agent measured against a 90% team target is a failure on paper and will behave like one.

Two things about what you do with the score. Bring it forward: the first review of a seasonal agent should land inside their first 48 hours of live work, because that is when feedback still changes habits rather than correcting them. And it should arrive as a conversation, not a number, which is the difference between turning QA data into coaching and simply reporting it. A score delivered during peak that produces no coaching is administrative overhead you cannot afford this month.

Calibrate before peak, and once during it

Reviewer agreement drifts fastest under time pressure, and peak applies three pressures at once. Reviewers are rushing. They are scoring contact reasons they have never scored before. And they are applying a modified scorecard that nobody has used yet.

Two sessions are enough.

  • Pre-peak, on the peak scorecard. There is no point calibrating the normal rubric in October if you plan to use a different one in November. Run the session on last peak’s conversations if you kept them, using this peak’s rules, and settle the arguments before they cost you anything.
  • Mid-peak, short. Thirty minutes, three conversations, week two or three. The goal is catching drift while it is still correctable, not being thorough. A rushed calibration that happens beats a proper one that gets cancelled.

One detail that gets missed. If you have kept the programme alive by deputising team leads to review, those people are the least calibrated reviewers you have and the least likely to be invited to a session designed for the regular QA team. Invite them first. The mechanics are the same as any other QA calibration session, the difference is only who is in the room and which rubric is on the screen.

The post-peak review almost nobody runs

Book it before peak begins, for the second week of the quiet period. Not later. Memories fade fast, and the seasonal cohort will have left, taking with them the only people who can tell you what the onboarding actually missed.

It should establish three things, and it is worth being strict that it establishes all three rather than turning into a general debrief.

1. What the relaxations actually cost

Compare reopen rate, escalation rate and customer dissatisfaction across the relaxed period against the equivalent normal-season window. This is the single most valuable output and almost nobody produces it. It also cuts both ways, which is what makes it interesting: if suspending personalisation for six weeks cost you nothing measurable, that criterion may not deserve the weight it carries for the other forty-six weeks either. Peak is an unplanned experiment on your own scorecard. Read the results.

2. Which failures were process, not people

Peak concentrates process defects until they are impossible to miss. If forty agents all failed the same criterion in the same week, that is a policy, a macro, a knowledge gap or a broken system, and coaching individuals for it wastes the coaching budget twice over. Sort failures by criterion and by contact reason before you ever sort them by agent, which is the same discipline as any other root cause analysis.

3. Who to rehire

You will run peak again. Keep a per-agent quality record for the seasonal cohort with a rehire recommendation attached, written while you still remember. Most teams lose this entirely and re-recruit from zero every year, then wonder why the ramp is the same length every time.

Then archive the peak scorecard, note the version change on the timeline, and put the standard rubric back with a date. The programme should restart on a named day, not drift back.

Quality debt, and why January is not a quiet month

Peak does not destroy quality, it borrows against it. Everything you chose not to look at still happened, habits formed in six unreviewed weeks are now habits, and the coaching that did not happen is still owed. Each line of that debt produces a recognisable symptom about four weeks later.

Debt taken on during peakWhat it looks like in JanuaryHow you repay it
Unreviewed conversations, especially from the seasonal cohort and new contact reasonsA blind spot in the season you most want to analyseRetrospectively review a small stratified batch in the first quiet week, purely for learning, with no scores published
Coaching that did not happenHabits formed under pressure that nobody correctedRestart one-to-ones before you restart reporting
Reviewer agreement driftThe first scores of the new year disagree with each otherOne full calibration session before the standard rubric goes back into use
Deferred process defects, the ones forty agents all hitThe same failure reappears in the spring, coached as forty individual issuesSort peak failures by criterion and contact reason before you sort them by agent

For retail, the repayment window collides with the returns wave. The National Retail Federation’s 2025 Retail Returns Landscape put total industry returns at 849.9 billion dollars, with an estimated 19.3% of online sales coming back, and returns contacts are harder than sales contacts. So book the repayment before peak starts: coaching restarts in week one of the quiet period, the calibration session and rubric switch-back in week two, and the retrospective batch and write-up in week three. Then send the leader you briefed in September the same one page, with actual numbers next to the forecast ones.

What changes when review coverage does not depend on reviewer headcount

Every decision above exists because of one constraint: manual review capacity comes out of the same pool of people as queue capacity, so the two compete directly. Sampling, budgets and allocation are all ways of rationing a fixed number of reviewer hours. Change that constraint and some of the decisions disappear while others get sharper.

What goes away is the rationing. When conversations are scored automatically, stratification stops being a budget exercise and becomes a reporting filter. Week-one agents and peak-specific contact reasons become cohorts you can look at whenever you want, rather than allocations you had to argue about in September. This is the argument behind the line Kaizo puts on its own service pages, that 100% coverage reveals trends 3% sampling never could, and peak is where that difference is most visible, because peak is when the population changes fastest.

What does not go away is the judgement. What to relax, what stays a gate, how to compare a seasonal cohort to a tenured one, whether a failure is a person or a process. No system decides those, and any vendor implying otherwise is selling you out of a decision you still have to make.

And one thing gets harder. Consistency cuts both ways. A miscalibrated criterion applied to a 3% sample distorts a handful of conversations and a good reviewer quietly compensates. Applied to every conversation it distorts all of them, in the same direction, during your highest-volume weeks. So the pre-peak calibration matters more when scoring is automated, not less.

Which means the thing worth checking before peak is not coverage, since most platforms now offer it, but whether you can show a score was right when someone disputes it. A seasonal agent challenging a deduction three days before Christmas deserves the specific exchange that triggered it, not a total. Kaizo’s Auto QA applies your own scorecard to conversations inside Zendesk and Salesforce Service Cloud and leaves the reasoning attached to the conversation, so a peak-period score can be traced back and argued from evidence rather than from memory. Reporting through Kaizo Insights then lets you hold the seasonal cohort and the tenured cohort apart without rebuilding the report every week.

Frequently asked questions

How do you manage high support ticket volume while maintaining quality support?

You decide in advance which parts of quality you will trade and which you will not, then you write it down. Relax consistency criteria like greeting, formatting and proactive personalisation. Never relax accuracy, verification and correct escalation, because errors in those three create a second contact at your most expensive moment. Shrink the review count to a number you can actually staff and spend it on new agents and new contact reasons rather than drawing it at random.

When should you start planning support quality for peak season?

Eight weeks before volume rises, and earlier if your ask involves spend or hiring. That is the last point at which budget exists, calendars have gaps and a substitute reviewer could still be trained. Two decisions belong in that window: whether you run quality reviews through peak at all, and the one-page brief telling leadership which numbers will move and what you are asking for.

Should you pause QA during peak season?

No, but you should resize it. Pausing removes visibility at the exact point your workforce, your contact reasons and your customers are least familiar, and it leaves a hole in the record for the season you will most want to study afterwards. Cut the review count to what one reviewer can genuinely deliver alongside queue time, cut the scorecard to the criteria that matter under load, and keep going.

How many conversations should you review during peak season?

Stop thinking in percentages and start with a number you can staff. A 3% sample of a normal month becomes an impossible workload when volume nearly triples, which is why programmes collapse rather than shrink. Fix an absolute review budget, then allocate it: roughly 40% to agents under 90 days, 30% to contact reasons that are new this season, 20% to business as usual so you can still see your baseline, and 10% to reopens and escalations.

How do you score seasonal customer service agents fairly?

Give them a shorter scorecard, not a softer one. Cut to five to seven criteria a competent person can meet in week one: correct answer, correct process, correct escalation, respectful tone, verification completed. Report the seasonal cohort separately from tenured agents so neither number misleads you, and compare each agent to their own cohort at the same tenure rather than to a team target built for people with two years on the job.

What should you tell leadership before quality drops at peak?

Send one page eight weeks out saying which numbers will move, in which direction, roughly how far, why in one repeatable sentence, when they come back, and the single thing you are asking for. Include a warning that scores produced under a modified peak scorecard are not comparable to the annual series.

What is quality debt and when do you pay it back?

Quality debt is what peak borrows rather than destroys: the conversations nobody reviewed, the coaching that did not happen, the reviewer agreement that drifted, and the process defects deferred because everyone was busy. It comes due about four weeks later. Repay it on booked dates rather than when things calm down: coaching first, calibration second, retrospective review third.

What is backlog in customer service?

Backlog is the set of conversations that have arrived and not yet been resolved, carried forward from one day to the next. It matters for quality because a backlogged thread tends to be longer, touched by more than one agent, and reopened more often, which makes it hard to attribute a score to any single person. Score the handoff explicitly, or pull multi-agent threads out of individual scoring and review them as a process problem instead.

When should you run the post-peak QA review?

Book it before peak starts, scheduled for the second week of the quiet period. Any later and memories have faded and the seasonal cohort has already left. It should establish three things: what the relaxed criteria actually cost you in reopens, escalations and dissatisfaction, which failures were process rather than people, and which seasonal agents you would rehire.

Go into peak knowing what your quality actually looks like

Bring last peak’s conversations and the scorecard you used. We will show you what your sample missed, where the seasonal cohort really sat against your tenured agents, and which of the criteria you relaxed turned out to matter.

Book a demo Explore Kaizo Insights

In Kaizo Assignments Assignments give manual reviewers an even, consistent distribution of work, targeted at the conversations where a human look actually changes something. See Assignments

On this page

See this on your own conversations

We will score a sample of your real tickets against your standards, so the example is yours.

Trusted by global support teams

  • Foot Locker
  • SteelSeries
  • Canva
  • GetYourGuide
  • Instacart