Call Sampling for Quality Assurance

Introduction

Most contact centers can't listen to every call. Managers pick a handful of recordings, score them against a rubric, and use that snapshot to judge how an agent — or an entire team — is performing.

That snapshot is smaller than most people assume. One ICMI and NICE study found that a large share of contact centers monitored just 1%-3% of interactions for quality. A tiny sample can still be useful, but only if it's selected the right way.

This guide is for contact centers, BPOs, answering services, and regulated teams that can't manually review every call but still need defensible QA results. You'll learn:

  • How call sampling actually works in QA programs
  • How to select calls more fairly across agents and queues
  • What biases and process gaps distort sample-based scores
  • When to move from small samples to broader interaction analysis

Key Takeaways

  • Call sampling estimates performance without full review; selection method decides what you can conclude.
  • Routine assurance sampling and targeted investigations answer different questions — don't blend their scores.
  • Stratify by queue, shift, call type, risk, and agent to avoid systematic blind spots.
  • Automated analysis extends coverage and surfaces red flags faster when human review can't keep pace.

What Is Call Sampling for Quality Assurance?

Call sampling pulls a subset of recorded interactions from a defined population, such as all inbound service calls handled last month, and scores them against a QA rubric within a set review period.

The result is meant to represent something larger than the calls themselves: an agent's typical performance, a queue's compliance rate, or a program's overall customer experience.

Teams use call sampling to:

  • Estimate how agents typically perform, not just how they performed on their best day
  • Flag coaching opportunities before they become patterns
  • Verify compliance without auditing every recorded interaction
  • Track customer experience trends over time

Sampling is not full interaction analysis, which examines every available call instead of a subset. It also differs from a targeted investigation (a review triggered by a specific complaint or escalation) and from live call monitoring, which is a review activity rather than a sampling design.

That distinction matters in practice. When RevClear Solutions' quality partner ECA reviewed calls manually, managers could get through less than 1% of total call volume each month.

The sample itself wasn't flawed. It simply couldn't reliably show how agents handled calls day to day. That gap between what a small sample shows and what is actually happening on the floor is the core problem call sampling exists to manage.

How Call Sampling Works and Why Contact Centers Use It

A defensible sampling process follows a consistent flow:

  1. Define the review objective: coaching, compliance, client reporting, or calibration
  2. Establish the eligible call population: which calls count, and which are excluded
  3. Select calls using a documented method
  4. Apply a consistent rubric across every selected interaction
  5. Validate findings through calibration or spot checks
  6. Use results for coaching, process changes, or compliance reporting

Six-step call sampling quality assurance process flow

Each step depends on inputs most teams already have: call recordings or transcripts, agent and queue metadata, review-period dates, quality criteria, and any documented exclusions.

Skip the documentation step and the program loses defensibility. Later, nobody can explain why certain calls were chosen and others were not.

Why Sampling Exists in the First Place

Contact centers sample because reviewer capacity is finite. High interaction volume, recurring performance checks, and the need to prioritize risky calls for human attention all push teams toward sampling instead of full review. Teams adopt it because full manual review cannot keep up with volume.

Six Sampling Approaches, Compared

Approach Best for Main limitation
Convenience Quick coaching example Nonrandom; easiest-to-access calls rarely represent the whole population
Fixed-number Scheduling coaching coverage Doesn't guarantee unbiased or precise selection
Random Routine performance baseline Small samples can miss rare but serious errors
Stratified random Comparing across queues, shifts, or languages Requires reporting group results separately
Systematic Simple, repeatable draws Periodic patterns (every 5th call) can quietly bias results
Risk-based/targeted Investigating complaints or flagged calls Deliberately skewed toward risk — not a routine error rate

Routine assurance sampling and targeted investigation solve different problems:

  • Random sample: estimates normal performance across the eligible population
  • Targeted review: investigates a complaint, escalation, or compliance concern
  • Reporting rule: keep targeted findings separate; do not fold them into the monthly average

There's No Universal Sample Size

Sample size depends on population size, the confidence level you need, and the margin of error you can tolerate. ASQ flags the same core inputs for any sample, QA included: desired confidence, allowable error, population characteristics, and event likelihood.

No single call count fits every agent, queue, or objective. High-risk compliance work needs tighter precision than a routine coaching check.

Building a Reliable Call Sampling Strategy

Start with purpose, not calls. Agent coaching, compliance assurance, client reporting, and calibration each require different selection and scoring choices. A sample built for coaching might tolerate more variance than one built for a client-facing compliance report.

Define the eligible population precisely. Decide whether the sample covers:

  • Inbound calls only, or outbound too
  • Transfers and callbacks
  • Abandoned interactions
  • Escalations
  • Only completed, recorded calls

Vague population definitions are where sampling programs fall apart.

Stratify Before You Select

Unstratified samples tend to overrepresent whatever's easiest to pull, usually daytime shifts on the main queue. Stratify across the dimensions that matter for your operation:

  • Agent, team, and tenure
  • Queue, call type, and customer segment
  • Shift, location, and language
  • Risk level and duration

Draw randomly within each stratum, and keep a record of which calls were selected, excluded, unavailable, or duplicated. That record is what makes the sample defensible later.

ICMI recommends reviewing high-value calls in addition to random calls, not instead of them. Routine random samples and supplemental targeted reviews serve different purposes. Don't let flagged or manager-requested calls inflate or deflate a routine score.

A short implementation checklist:

  • Document eligibility rules and exclusions in writing
  • Use one consistent scorecard across reviewers
  • Calibrate reviewers periodically against a reference score
  • Audit sample composition to check queue, shift, and language coverage
  • Revisit the strategy when call volume or business risk changes

Consistency beats complexity on the scorecard itself. ECA's rubric covered greetings, hold handling, caller verification, message accuracy, tone, and pacing: a compact list, applied the same way every time, rather than a sprawling scorecard nobody could score twice the same way.

Key Factors, Limitations, and Common Misconceptions

Several operational factors shape what a feasible sample size actually looks like:

  • Call volume and agent workload
  • Reviewer capacity and scorecard length
  • Review frequency
  • Confidence level the program needs

More reviewers or shorter scorecards can raise feasible sample size, but they don't fix a biased selection method.

Common distortions worth watching for:

  • Missing recordings that quietly exclude certain call types
  • Transcription errors, especially in accented speech or non-English calls
  • Uneven shift coverage that favors daytime agents
  • Inconsistent reviewer availability skewing which calls get scored

"Bigger Sample" Doesn't Mean "Representative Sample"

A large convenience sample can still exclude night shifts, difficult queues, or unusual call types entirely. Size alone doesn't fix a biased selection frame. A well-stratified sample of 50 calls can be more informative than an unstratified sample of 500.

Comparison of sample size and representative call quality

A related misconception: a fixed number of calls per agent sounds fair. Equal quotas still ignore real differences in call volume, role, risk exposure, tenure, and interaction complexity. An agent handling twice the call volume of a peer, or working a higher-risk queue, may need a different review depth entirely.

Mixing review types: Escalated calls, randomly selected calls, and manager-requested calls scored together produce a number nobody can interpret. Was last month's dip in scores a real performance issue, or did a cluster of escalations get folded into the routine sample? Label each review type separately, and report them separately too.

Clean separation only helps if the results drive change. Findings should feed back into updated coaching plans, revised training, workflow fixes, or scorecard changes, not sit as a monthly metric nobody revisits.

When Call Sampling May Not Be Enough

Manual sampling starts to strain in specific situations:

  • High-volume contact centers where even a "large" sample covers a tiny fraction of calls
  • Multi-site operations trying to standardize scoring across locations
  • BPO programs juggling client-specific scorecards across accounts
  • Regulated sales or service calls where a single missed disclosure carries real risk
  • Teams that need faster red-flag detection than a monthly sample review allows

Gartner frames evaluating 100% of service interactions as the goal automated QA is working toward, not a claim that every organization already scores every call that way.

A hybrid model is the practical middle ground: automated analysis scores a broader set of interactions, then routes the highest-priority calls and representative examples to human reviewers for validation, coaching, and calibration.

This is the gap platforms like EmberQA are built to close. When RevClear Solutions' partner ECA moved from manual sampling to EmberQA, coverage went from under 1% of calls to 100%. Every call was transcribed, scored against ECA's rubric, and returned with an explanation for the score.

Manual versus automated call quality assurance coverage comparison

Managers stopped guessing which calls were worth a closer listen; the system surfaced them. That does not replace rubric design or human judgment. Someone still has to build the scorecard and calibrate the scoring, but automated coverage removes the ceiling manual sampling puts on review volume.

Conclusion

Call sampling works, but only as far as its design holds up. A clear objective, a defensible selection process, a consistent rubric, and transparent reporting are what separate a useful sample from a number that just sounds precise.

Routine random or stratified samples and targeted investigations answer different questions. Keep them separate, and report them separately.

The right move isn't always "sample more." Sometimes it's tightening the sampling design you already have. Other times, when volume, risk, or coaching needs outgrow what a subset can reliably show, broader automated analysis is the more honest way to see what happens on every call.

Frequently Asked Questions

What is sampling?

Sampling means selecting a subset of a larger population for analysis instead of examining everything. In QA, that means reviewing a portion of recorded calls to draw conclusions about a larger set of interactions.

What are examples of sampling?

Common QA sampling methods include random sampling, stratified sampling (by queue, shift, or language), systematic sampling, fixed-number sampling, and targeted or risk-based sampling for flagged calls.

How many calls should be sampled for quality assurance?

There's no universal number. Size depends on call volume, review objective, risk level, required confidence, and reviewer capacity. Routine coaching samples and compliance-driven samples usually need very different sizes.

What is the best way to select calls for QA review?

Start with documented eligibility criteria, stratify by relevant factors like queue and risk, then select randomly within each group. Handle escalated or high-risk calls as a separate review stream, not part of the routine sample.

What is the difference between random and targeted call sampling?

Random sampling supports a view of typical performance across a population. Targeted sampling investigates known risks or specific events and shouldn't be blended into routine QA scores.

Can AI replace call sampling for quality assurance?

No. AI can score every interaction and flag risks sampling would miss, but it doesn't replace human judgment. Reviewers still own rubric design, calibration, and final calls on edge cases.