Skip to content
    August 16, 2026| Top Floor Team| 9 min read

    How Audit Sampling Works: How Many Items Will They Pull?

    Auditors do not check every ticket. They define the population of a control's occurrences, draw a sample from it, and test those items, and no standard tells them how many to pull. PCAOB AS 2315, the audit sampling standard, sets no minimum sizes at all; it lists the factors that drive the decision (risk assessments, tolerable misstatement, the expected size and frequency of deviations, and the characteristics of the population) and leaves the number to professional judgment. The AICPA's own sampling guidance for practitioners, AU-C section 530, works the same way.

    What actually decides your outcome is not the sample size. It is whether the examiner can rely on the population you handed over, because AS 2315 requires them to "determine that the population from which he draws the sample is appropriate for the specific audit objective", and an incomplete population makes everything drawn from it worthless.

    That is the short version. The rest is how sizes really get set, how a deviation gets read, and what to do about completeness before anyone asks.

    Key takeaways

    • No standard fixes a sample size. PCAOB AS 2315 sets no minimums and lists the factors; AICPA AU-C 530 works the same way, leaving the number to professional judgment.
    • Three things move the number, and none of them is your company's size: how often the control runs, how much the examiner needs to rely on it, and what they expect to find.
    • A failed item is a deviation, not automatically an exception. What decides its fate is whether it is isolated or systematic, what the rate implies about the untested remainder, and whether anything else caught it.
    • Completeness is what actually fails companies. An incomplete population makes everything drawn from it worthless, and examiners reconcile your list against an independent source.
    • Run your own miniature fieldwork at roughly month three and month nine, against the same populations the examiner will use. It is the only leverage you have, because nothing lets you retroactively perform a review you skipped.

    What drives the number

    Three things move a sample size, and none of them is your company's size.

    How often the control runs. A control that operates once a year has a population of one, and the examiner tests that one. A control that operates every time an engineer deploys might have a population of two thousand. Frequency drives population, and population influences the sample, though not proportionally: sample sizes grow far more slowly than populations do.

    How much the examiner needs to rely on it. A control that is the sole thing standing between an outsider and customer data will be tested harder than one of four overlapping controls in the same area. This is why the same control can be sampled differently at two companies.

    What the examiner expects to find. AS 2315 is explicit that expected deviations feed the size. If your first-cycle access review evidence came back messy, expect a larger sample in the areas that looked shaky, not a smaller one.

    Firms publish their own internal sizing tables and most will share the shape of theirs if asked. That is the right way to get this number: ask your examiner, in writing, before the observation window closes, "for each recurring control, what population will you request and roughly how many items do you expect to test?" A firm that answers has a methodology. Do not take a number off someone else's blog and plan your evidence around it, because the number you need is the one your examiner will use.

    Deviations, and how one gets read

    When a sampled item fails, that is a deviation, and it does not automatically become an exception in your report. What happens next depends on how the examiner reads it.

    The first question is whether it is isolated or systematic. One late deprovisioning out of twenty-five, where the same control worked every other time, reads as a lapse. Three late deprovisionings all involving the same system that sits outside your identity provider is a pattern, and a pattern says the control has a hole rather than a bad day.

    The second question is what the deviation rate implies about the untested remainder. This is the whole logic of sampling: the sample stands in for the population, so a deviation rate in the sample is evidence about the rate in the population you did not test. Two failures out of twenty-five is an eight percent deviation rate against a control you described as always operating, and the examiner has to consider what eight percent means across the whole period.

    The third question is whether anything else caught it. A monthly dormant-account review that found and disabled the orphaned account, with logs showing it was never used after termination, changes the character of the finding. It does not erase it, but it is exactly the sort of compensating evidence that keeps a deviation from escalating. We walk through how that escalation actually works, from exception to qualified opinion, in can you fail a SOC 2 audit.

    Completeness is what fails companies

    Here is the failure mode we see most often, and it has nothing to do with how many items get pulled.

    You hand over a list of the year's terminations. The examiner does not simply sample it. They reconcile it against an independent source, usually payroll or the HR system of record, and they find four people on the payroll list who are not on yours. Three were contractors, and one was an employee who transferred to a subsidiary. Now the problem is not that your offboarding control failed. The problem is that your population was wrong, so nothing sampled from it supports any conclusion, and the examiner has to either expand testing or conclude they could not rely on the control at all.

    The same trap sits under every population an examination touches:

    • Terminations reconciled against payroll rather than against IT's ticket queue
    • Production changes reconciled against the deploy log rather than against the ticket system
    • Users reconciled against the identity provider and against each system that sits outside it
    • New hires reconciled against payroll, which is where the contractor gap usually shows up
    • Incidents reconciled against the alerting system and the on-call log, not just against the incidents somebody remembered to write up

    The fix costs almost nothing if you do it early and costs a re-test cycle if you do it late. For every recurring control, name the authoritative system today, produce the population from it as a dated system export, and reconcile it once a quarter against a second source. Do that and you will discover your gaps in month three, when they can still be fixed, rather than in month eleven, when they cannot.

    What you can do while the window is open

    Sampling rewards a specific habit: mid-period self-testing, done against the same populations the examiner will use.

    Twice during the observation window, ideally around month three and month nine, run your own miniature version of fieldwork. Pull the full population for each recurring control. Take ten items at random. Ask for the evidence exactly as an examiner would, and see what comes back. It takes a few hours and it tells you two things you cannot learn any other way: whether the evidence exists, and whether it can be found by someone who was not there when it was created.

    A control that failed in month two and ran cleanly for ten months reads completely differently from one that was still broken at fieldwork. Nothing in the standards lets you retroactively perform a review you skipped, so the only leverage you have is time, and time only helps if you spend it knowing.

    When sampling does not apply

    Two cases worth knowing, because people over-prepare for them.

    Controls with a population of one are tested directly, not sampled. Your annual risk assessment, your annual policy review, your annual penetration test, your disaster recovery test: there is one occurrence, and the examiner looks at it. There is nowhere to hide and no sampling risk, which is why these are the easiest items to get right and the most embarrassing to miss.

    Fully automated controls with reliable configuration evidence are sometimes tested through the configuration plus a small confirmation sample rather than through a large sample of occurrences. If the identity provider enforces MFA for every login and the examiner can see the policy and its change history, they may not need to inspect a hundred logins. This is the mechanism behind the real benefit of automation: not fewer controls, but cheaper evidence for the controls a machine enforces.

    Where Top Floor fits

    Most of the value here is unglamorous: knowing which system owns each population, producing the exports, running the mid-period self-test, and having the argument with the examiner about whether a deviation is isolated. That is the day job of audit readiness and facilitation work, and when we run the program itself it is the part that quietly determines how the report reads.

    If you have an operations-minded person who will own the compliance calendar and enjoys reconciling lists, you do not need to buy this. Two recurring tickets and one quarterly reconciliation cover most of it. What that person needs is the authority to chase engineers, which is a different thing from being assigned the task.

    Frequently asked questions

    How many samples will a SOC 2 auditor pull?

    There is no mandated number. PCAOB AS 2315 and AICPA AU-C 530 both leave sample size to professional judgment, driven by the control's frequency, how much the examiner needs to rely on it, and how many deviations they expect. A yearly control has a population of one and is tested directly; a daily control has a large population and a sample that grows far more slowly than the population does. Ask your own examiner for their expected populations and sizes in writing before the observation window closes, rather than planning against a number from someone else's methodology.

    What happens if one sampled item fails?

    It becomes a deviation, and the examiner weighs three things: whether it is isolated or part of a pattern, what the deviation rate implies about the items they did not test, and whether any other control caught it during the period. An isolated lapse with compensating evidence usually ends as an exception noted in the report rather than a change to the opinion. A cluster in one criterion, with nothing else catching it, is what pushes an opinion toward qualification.

    Why does the auditor reconcile our list against payroll?

    To test completeness. The sample only supports a conclusion if the population it came from includes every occurrence, so examiners compare the list you provide against an independent source: payroll for terminations and hires, the deploy log for changes, the identity provider for users. An incomplete population is worse than a failed sample, because it means no conclusion can be drawn about the control at all.

    Can we test our own samples before the audit?

    Yes, and it is the highest-value hour of preparation available. Twice during the observation window, pull the full population for each recurring control, select ten items at random, and request the evidence exactly as an examiner would. You will find out whether the evidence exists and whether someone who was not present can locate it. Do it in month three and month nine, while a discovered gap is still fixable.

    Share Share on LinkedIn

    Need help with your compliance program?

    Our team of senior practitioners can help you navigate complex compliance requirements and build a security program that holds up under scrutiny.

    Schedule a Free Consultation

    Get insights like this in your inbox

    Practical compliance and security guidance for teams preparing for their next audit. No spam, unsubscribe anytime.

    Ask to be added to our mailing list for practical compliance and security guidance. We add you by hand, we confirm before sending anything, and we never share your address.