Learn · SOC Reports
SOC 2 Sampling, Populations
and Deviations
Almost every argument in a SOC 2 Type 2 fieldwork cycle is really an argument about sampling — what the population was, whether it was complete, how many items the auditor picked, and what happens when one of them fails. The mechanics are rarely written down. Here they are.
The order matters more than the numbers. Population first, completeness second, selection third. A sample drawn from an incomplete population is not weak evidence — it is evidence about the wrong thing, and no sample size repairs it.
Plain-English explainer · AT-C 205 examination · Last reviewed August 2026
In a SOC 2 Type 2, the auditor does not inspect every occurrence of every control. They define a population — every occurrence of that control inside the observation period — prove it is complete and accurate, select a subset from it, and form a conclusion about the whole population from what those items show. Sample sizes are set by the CPA firm’s own methodology, not by the AICPA, and are driven by how often the control operates, how long the period is, and how many deviations the auditor expects to find. There is no AICPA table of required SOC 2 sample sizes. The attestation sections that govern the examination — AT-C 105 and AT-C 205, as recodified by SSAE 21 — set no sample sizes at all, and nothing in them says a monthly control gets two items. The ranges everybody quotes are firm methodology, anchored to guidance written for financial statement audits and adopted into SOC practice by convention. Knowing which parts are rules and which are convention is the difference between negotiating usefully with your auditor and arguing about the wrong thing.
First principle
The population is defined by the control text
A population is the complete set of items to which a control was applied during the observation period, and what counts as an item is decided by the wording of the control — which is why control language is a sampling problem before it is a documentation problem. Two shapes recur. A scheduled control operates on a calendar, so its population is its own occurrences: a quarterly review has a population of four. An event-driven control operates whenever something happens, so its population is the events — every joiner, every leaver, every deployment, every alert above the threshold.
| Control | The population is | What quietly breaks it |
|---|---|---|
| Quarterly user access reviewCC6.3 | The four scheduled reviews in the period — a calendar control, so the population is its occurrences, not the users reviewed. | A review performed but never evidenced; a quarter skipped while the reviewer was on leave and never logged. |
| Access granted to new joinersCC6.2 | Every person who received system access during the period — an event-driven control, so the population is the events. | Contractors provisioned outside the HRIS; service accounts created by engineers; a mid-period HR migration stranding joiners in the old tenant. |
| Access revoked on terminationCC6.3 | Every leaver in the period — voluntary and involuntary, employee and contractor. | An export filtered to "employees" that silently drops contractors; a leaver whose record was updated after the export ran. |
| Peer review of production changesCC8.1 | Every change deployed to production in the period, across every deployment path — not only the primary pipeline. | Emergency hotfixes applied by hand; infrastructure changes applied from a laptop; a second repository nobody raised in scoping. |
| Alert triage and investigationCC7.2 · CC7.3 | Every alert meeting the severity threshold the control names — vague control text produces an untestable population. | A control that says "significant alerts are investigated" without defining significant; alerts auto-closed by a tuning rule and excluded from the export. |
| Vendor security review before onboardingCC9.2 | Every new vendor onboarded in the period that meets the criticality definition in the control. | Vendors bought on a corporate card outside procurement; renewals treated as "not new" when the control says they are in scope. |
The right-hand column is where the pain lives, and it is always the same failure: a population reflecting the happy path through one system rather than every path. The system boundary described under the AICPA description criteria (DC section 200) settles whether contractors, service accounts, a second region or a legacy product belong in the population. If the description says they are in the system, they are in the population, and an export that omits them is incomplete however clean the sampled items look.
Completeness and accuracy
Why the population is tested before the sample
A sample supports a conclusion about the population it was drawn from and nothing else. If forty joiners were provisioned during the period but the listing shows thirty-two, testing five of the thirty-two says something about those thirty-two and nothing about the eight that were never on the list — and the eight are exactly where an unapproved or over-privileged account is most likely to sit. An incomplete population does not weaken the sample; it invalidates the inference the sample was meant to support.
The listing is itself evidence, and you produced it. The attestation standards require the practitioner to evaluate whether information used as evidence is sufficiently reliable — obtaining evidence about its accuracy and completeness, and judging whether it is precise and detailed enough for the conclusion being drawn. SOC practice calls this information produced by the entity, or IPE, and it is not credible merely because a computer generated it. A system-generated report is only as reliable as the query behind it, the parameters someone typed, and whatever happened to the file between export and delivery.
Completeness — is anything missing?
Tested by understanding how the report is built and what its filters exclude, reconciling record counts to an independent view of the same data, and checking that known items appear. If you know two people left in July, both belong in the leaver listing. If the deployment log shows 621 releases and the export shows 612, the nine are the finding — before any sampling begins.
Accuracy — are the fields right?
Tested by tracing individual records back to the underlying system: does this termination date match the HR record, does this ticket carry that approver, does this role match the identity provider. Completeness and accuracy are separate tests, and a listing can pass one and fail the other.
Where the auditor cannot get comfortable with a listing they have three moves: perform other procedures, source the information elsewhere, or decline to rely on it. The third is expensive — it usually means the control cannot be tested as designed, a far worse outcome than a deviation. Compliance automation platforms do not change this: a vendor dashboard listing your changes or users is IPE too, and the auditor still has to establish how the platform derived the list.
Extent of testing
How many items and who decides
Extent is a judgement the CPA firm makes, driven by the frequency of the control, the size and homogeneity of the population, the risk it addresses, and the deviation rate the auditor expects before starting. Firms encode that judgement in a methodology and apply it consistently, which is why the numbers look like a table even though no standard sets them.
Label this correctly. The table below is market practice, not standard. It reflects ranges in common use across SOC 1 and SOC 2 engagements, tracing back to the AICPA’s audit sampling guidance for tests of controls in financial statement audits. Neither the AICPA nor any other body mandates sample sizes for SOC engagements. Your auditor’s numbers may sit anywhere in these ranges or outside them, and the justification is a reasoned methodology — not the table. The three sample columns are the same kind of convention: they show how a firm’s starting extent moves once deviations are expected or found, not a schedule anyone is obliged to follow.
| Control frequency | Population (12 months) | Sample at 0 expected deviations | at 1 | at 2 | Note |
|---|---|---|---|---|---|
| Annual | 1 | 1 | 1 | 1 | The single occurrence is the whole population, so no extension exists. One deviation is a 100% rate. |
| Quarterly | 4 | 2 | 4 | 4 | Extending means testing all four. Some firms start at four where the control is the only one addressing a criterion. |
| Monthly | 12 | 2 | 4–5 | all 12 | The most disputed row. Two is defensible at zero expected deviations; one miss usually takes it to four or five. |
| Semi-monthly | 24 | 3–8 | 10–12 | all 24 | Wide starting range — firms weight the expected deviation rate differently at this size. |
| Weekly | 52 | 5–9 | 12–15 | 20+ | Selection is normally spread across the period, not clustered in one quarter. |
| Daily | ~250 | 25 | 40 | 60 | Business days over twelve months. Twenty-five is the usual starting point. |
| Many times per day | 250+ | 25 | 40 | 60 | Once the population is large it stops driving the number; the expected deviation rate takes over. |
| Event-driven, small population | e.g. 18 | often all 18 | all 18 | all 18 | Below roughly 25 items sampling stops saving effort, so there is nothing left to extend into. |
All of this is Type 2 material. A Type 1 examination reports on the suitability of the design and the implementation of controls as of a specified date rather than on operating effectiveness across a period, so it involves no sampling of occurrences at all — the auditor typically inspects a single instance to confirm the control was implemented. That is exactly why a Type 1 tells a reader nothing about whether the control kept working.
Two adjustments matter more than the rows. First, the population scales to the observation period: a monthly control in a three-month Type 2 has a population of three, not twelve, which is one more reason a short first observation period buys less assurance, quite apart from how it ages. Second, expected deviations push sizes up before testing starts. The first column above assumes none are expected; where readiness work or a prior period suggests otherwise, the starting sample moves along the row before a single item is selected, because a sample sized for zero deviations cannot absorb one.
Interim testing and the roll-forward
Most twelve-month Type 2s are tested twice
The table reads as though fieldwork happens once, after period end. On a twelve-month examination it usually does not. The auditor commonly performs interim fieldwork with two or three months still to run, then returns after period end for a roll-forward covering the remainder. That splits every population in two and changes the extent of testing for every control on the list.
What the split looks like with numbers
Observation period 1 April 2025 to 31 March 2026. The auditor performs interim fieldwork in December 2025 covering April to November: a monthly control has an interim population of eight and a sample of two. They return in April 2026 for the roll-forward covering December to March: population of four, sample of one or two. The full-period population is still twelve — it is the testing that arrives in two instalments, and the instalments are added together when the deviation rate is evaluated, not judged separately.
What the auditor must establish at roll-forward
Three things, and they are why the roll-forward is not a formality. That the control did not change between the two phases — a new approver, a new threshold or a new tool makes it a different control and two sub-populations in substance. That the population export for the roll-forward period was produced on the same basis as the interim one: same report, same filters, same source. And that any deviation found at interim was addressed rather than left running into the second phase, where it stops being an isolated item and starts looking like a pattern.
The practical consequence for you
Population exports get requested twice, and they have to match. An export whose filters, columns or underlying report changed between December and April is a completeness failure discovered at the worst possible moment — after period end, with no room left to remediate inside the period. Save the report definition at interim, record who ran it and how, and produce the second export from that saved definition rather than rebuilding it from memory.
Selection
How the twenty-five items get picked
Extent answers how many. Selection answers which — and it is the half of sampling almost nobody writes down, even though it decides whether the phrase “statistically valid” is available to describe the result at all. Four methods are in use. Only the first supports projecting the sample deviation rate onto the population as a statistical estimate; the other three are non-statistical, which is not a criticism, only a limit on what may be claimed from them.
Random
Every item has an equal, known chance of being selected
A seeded random number generator, an audit tool’s selection routine, or RAND over the row numbers of the population export. Applied to the 621 deployments, it returns 25 row numbers with no human eye on the list. This is the only method that supports statistical extrapolation — projecting the observed rate onto the population with a stated confidence — and so the only one that makes a “statistically valid” claim defensible.
Systematic
Every n-th item after a random start
The interval is the population divided by the sample size: 621 / 25 gives an interval of 25, so a random start at row 7 selects rows 7, 32, 57, 82 and onward. Fast, reproducible, and easy for a client to verify afterwards. It fails when the listing is sorted in a way that correlates with the interval — deployments ordered by service, joiners ordered by department — because a fixed stride through a sorted list can systematically miss or over-weight a group.
Haphazard
Selected without a randomising device and without conscious bias
The auditor works through the listing and picks items without a rule and without steering toward or away from anything in particular. This is the most common method on SOC engagements, and it is explicitly not statistical: unbiased in intent, but with no measurable probability of selection, so nothing can be projected from it with a confidence level attached.
Judgemental
Deliberately weighted toward the higher-risk items
The auditor targets what matters most — every privileged-role grant, every change deployed outside business hours, every vendor above a criticality threshold — sometimes stratifying the population and selecting from each stratum separately. Non-statistical by construction, because the weighting is the whole point. It produces stronger evidence about the risky items and a weaker basis for generalising to the ordinary ones.
Three things about how this runs in practice. The auditor spreads the selection across the observation period rather than clustering it in one month, because twenty-five items all drawn from May would evidence May. You then receive a written selection list — item identifiers with a response due date, not descriptions — and that list is the auditor’s work product: your job is to return the evidence for those identifiers, not to comment on which ones they are. And it is entirely fair to ask which method was used, because “statistically valid” and “haphazard” cannot both be true of the same selection.
Control type
Why one screenshot is sometimes enough
A manual control depends on a person doing something correctly each time, and people vary, so testing has to span the period — the evidence for March says nothing about September. An automated or configured control differs in kind: the system applies the same rule to every instance because the rule lives in a configuration — an enforced password policy, mandatory multi-factor authentication, a session timeout, a storage bucket blocking public access, a branch protection rule requiring an approving review.
For those, inspecting the configuration once can evidence how it behaved across the period — but the sufficiency is borrowed, not inherent. It rests on the auditor being able to rely on the general IT controls that stop the configuration changing without authorisation: change management under CC8.1, and restricted administrative access under CC6.1 and CC6.3.
When a single test holds
The configuration is genuinely system-enforced, administrative access to change it is restricted and tested, change management over the platform is operating effectively, and a configuration-change history for the period shows either no changes or only authorised ones.
When the auditor tests more
Administrative access is broad, the platform sits outside change management, the configuration was altered mid-period, or no change history exists. The auditor then typically inspects the configuration at more than one point — commonly near both ends of the period — or tests a sample of the instances it should have governed.
The common misreading
Teams see "tested once" in a prior report and assume every technical control gets a single screenshot. It is the ITGC reliance that earns the single test, not the involvement of a tool. A control performed in a tool by a human — reviewing an alert queue, approving a ticket — is still manual and is sampled across the period like any other.
So automation reduces the evidence burden as well as the chance of failure — but only if you also bring the platform holding the configuration inside change management and access control. Automation without those two is the worst of both worlds: a control nobody performs manually and an auditor who cannot rely on the configuration.
Deviations
A deviation is not yet an exception
A deviation is a departure from the control as described: the sampled item shows it did not operate the way the control language says it does. An exception is what appears in Section 4 of the finished report. Every exception began as a deviation; not every deviation ends as one, and even an exception does not by itself modify the opinion.
Between the two sits an evaluation. The auditor investigates the nature and cause of what they found, considers whether it is confined to something identifiable or symptomatic of the control generally, and assesses the effect on the conclusion. Sometimes a deviation dissolves — the item was outside the population, or the evidence sat somewhere the requester had not looked.
A useful discipline from audit sampling guidance: where deviations share a common feature — the same person, location, fortnight or deployment path — the auditor may identify every item in the population carrying that feature and extend procedures to all of them, converting a vague worry into a bounded sub-population with a measured rate. United States practice is deliberately more conservative here than the international equivalent: ISA 530, Audit Sampling, recognises an anomaly — a deviation the auditor can demonstrate is not representative of the population — and permits it to be excluded from the projection, while the AICPA’s audit sampling standard for financial statement audits, AU-C section 530, Audit Sampling, deliberately omits the concept. AT-C 205 introduces no equivalent of its own, so US SOC practice inherits the more conservative treatment.
| Control | Pop. | Sample | Devs | Rate | How it reads |
|---|---|---|---|---|---|
| Annual DR test | 1 | 1 | 1 | 100% | The control did not operate. Nothing remains to evaluate about representativeness. |
| Quarterly access review | 4 | 2 | 1 | 50% | Half the tested occurrences failed. Testing all four is usually the only way forward. |
| Monthly reconciliation | 12 | 3 | 1 | 33% | Extension is cheap — nine untested items remain, and most firms will take them all. |
| Production changes | 621 | 25 | 1 | 4% | Low enough that asking whether the deviation is confined to a sub-population is meaningful. |
That arithmetic is the most under-appreciated fact about SOC 2 testing. The deviation rate an auditor will tolerate does not relax because a population is small — but the smallest rate it is possible to observe grows enormously. Low-frequency, high-consequence controls — the annual disaster recovery test under A1.3, the annual risk assessment, the quarterly access review — carry almost no tolerance in practice, and they are precisely the ones organisations let slip a cycle.
Extending the sample is the usual next step on a large population. The 25 → 40 → 60 pattern is market practice in exactly the way the sample-size table is, not a rule: a common firm methodology runs 25 items, extends to 40 after one deviation and to 60 after a second, then stops extending — at which point most methodologies conclude the control did not operate effectively for the period. Extension buys information about whether the failure was confined; it never removes the deviation, which stays in the numerator. On a small population extension is barely available at all.
“Can we swap that item out and give you a different month?”
— the request that ends an engagement’s credibility. Selection belongs to the auditor; a sample the client chose is not evidence, and substituting an item after seeing it fail turns a manageable exception into a scope problem.
Worked example
One deviation, followed through
An illustrative engagement — the shape is drawn from how these evaluations run, not from any particular client. Type 2 examination, observation period 1 April to 30 September 2025. Control CHG-02: all changes to production are peer-reviewed and approved in the change record before deployment. Mapped to CC8.1.
Population
The first export, from the CI/CD platform, listed 612 production deployments. Reconciled to the production deployment log, the true figure was 621 — the nine-item gap was emergency hotfixes applied outside the pipeline. They were added before selection. Had the gap surfaced after testing, every tested item would have needed re-evaluating.
Selection
Population above 250 and no deviations expected, so a sample of 25 was drawn across the six months by the auditor — not by management.
Deviation
One item failed. Change CHG-4471, deployed Saturday 12 July 2025, carried an approval recorded 14 minutes after the deployment timestamp rather than before it. The approval existed; the sequence the control requires did not.
Nature and cause
An on-call engineer had used the emergency deployment path for a non-emergency fix during an unrelated incident. The cause was a process gate that did not require an incident reference — not an absent review culture, a distinction that changes what remediation is credible.
Common feature, extended
The deviation came through the emergency path, so the auditor tested all nine emergency-path changes and found one more with the same sequencing failure: two deviations in a bounded sub-population of nine. CHG-4471 was itself one of the original 25, which left 24 standard-pipeline items in that selection; those 24 were extended by 16 further pipeline changes to 40, with no additional deviations. The emergency-path items were evaluated separately as the bounded sub-population of nine, so the failed item is counted once, in the sub-population it belongs to.
Evaluation and reporting
The auditor concluded the failure was confined to the emergency path, that the path itself required documented review within one business day and both instances met that, and that peer review over the 40 standard-pipeline changes tested operated without exception. The exception was described in Section 4 and the opinion was unmodified. Management’s response in Section 5 recorded a gate added on 6 October 2025 requiring an incident identifier before the emergency path can be used.
Change the numbers and the conclusion changes with them. Six deviations in the nine emergency-path changes would have made the sub-population the story, and a reasonable auditor would have concluded the criterion was not met for that path. The evaluation is genuinely an evaluation — which is why “how many exceptions are allowed?” has no answer, and why opinions and exceptions have to be read together rather than counted.
What that becomes in Section 4
Readers almost never see the language a service auditor actually writes, which is why Section 4 looks impenetrable the first time. This is the anonymised shape of the row for CHG-02 — the same three parts every Section 4 row has.
Control
All changes to production are peer-reviewed and approved in the change record before deployment.
Test performed by the service auditor
Inspected the listing of production deployments for the period 1 April 2025 to 30 September 2025, reconciled the listing to the production deployment log, and determined the population to be complete. For a sample of 25 deployments selected from the population of 621, inspected the associated change record for evidence of documented peer review and approval recorded prior to the deployment timestamp. Extended testing of standard-pipeline deployments to 40 items and inspected all 9 deployments effected through the emergency deployment path.
Results of tests
Exception noted. For 2 of the 9 emergency-path deployments selected, approval was recorded after the deployment timestamp. No exceptions were noted for the 40 deployments effected through the standard pipeline.
Four things in that middle paragraph are what to look for in any test step: reconciled, the word population, the explicit sample of 25 from 621, and prior to the deployment timestamp. Between them they tell a reader that completeness was tested, what the population was, how far the testing went, and what the auditor actually looked at inside each item. A test step reading only “Inspected evidence of approval” says nothing about extent — and a Section 4 written that way throughout is the single most useful thing to notice when comparing two reports.
Nil populations
When the population is genuinely zero
Some event-driven controls never fire. No security incidents were declared under CC7.3 and CC7.4, no employees were terminated, no emergency changes were raised. This is normal — particularly for small teams and short periods — and it is reportable, but it demands two things teams routinely skip. First, the zero has to be evidenced: a query returning zero rows with its parameters and date range visible, not an assertion that nothing happened, because a zero produced by a wrong filter looks identical to a real one. Second, a nil population must be distinguished from a control that did not run. They look alike on a spreadsheet and are not alike at all.
Event-driven, no events — a nil population
Nothing triggered the control. Section 4 typically records that no instances occurred during the period and that no operating-effectiveness testing was therefore performed; some auditors add a short passage to the auditor’s report identifying controls that did not operate. It is not an exception. It also carries no assurance — a reader learns nothing about your incident response from a report where incident response never operated.
Scheduled, did not happen — a deviation
A quarterly access review with three reviews instead of four; an annual recovery test that slipped past the period end; a risk assessment nobody ran. The population is not zero — it is the scheduled occurrences, and one is missing. That is a deviation with a 100% or 50% sample rate attached, and describing it as “not applicable” changes nothing about what the auditor concludes.
The forward-looking move on a nil population is to consider whether the control can be exercised deliberately: a tabletop incident simulation, a recovery test, a revocation drill on a test account — performed and evidenced inside the period — turns a silent control into a tested one. Whether that satisfies the control as written depends on the control language, so it is a readiness conversation with your auditor, not a unilateral decision in month five.
Evidence
What the CPA firm actually requests
Two separate requests arrive, and conflating them causes most of the rework. The first establishes the population; the second tests the selected items. Evidence that fails the first never reaches the second.
| What is requested | What satisfies it | What gets sent back |
|---|---|---|
| The population export | A direct, unedited export from the source system covering exactly the observation period. | A spreadsheet assembled by hand, a copy-paste into a new tab, or a file with rows deleted "because they were not relevant". |
| Report parameters | Report name, source system, filters applied, date range, who ran it, and the date and time generated — visible on screen, not described in an email. | An export with no evidence of the filter used. If the auditor cannot see the parameters, they cannot conclude it is the whole population. |
| The query or report logic | For anything custom: the SQL, the API call, or the saved report definition — so the auditor can see what it includes and, more importantly, excludes. | A number with no derivation. "The system says 621" is a claim, not evidence about how 621 was arrived at. |
| Record counts and a tie-out | A count reconciling the export to an independent view of the same data — a dashboard total, a second system, a deployment log. | Counts that disagree with no explanation. An unexplained variance is a completeness failure, not a rounding issue. |
| Screenshots with context | Full-window captures showing the URL or server name, the logged-in user, the menu path and the system clock. | Crops with the chrome removed, images with no timestamp, or a filtered view with the filter bar cut off. |
| The sampled items | For each selected identifier, the artefact showing the control operated on that item — and the artefact differs by control type, which is where almost all of the rework happens. Six common ones are set out in the second table below. | Evidence recreated after the fact. A review re-performed during fieldwork proves the control can operate, not that it did. |
The second request is where evidence is most often rejected, because the artefact that satisfies a control is specific to that control and rarely what the team assumed. Six that come up in almost every engagement:
| Sampled item | The artefact that satisfies it | What gets sent back |
|---|---|---|
| Termination revocation | The HRIS termination record showing the effective date, plus the identity-provider or SSO deprovisioning event log entry carrying its own timestamp — so the elapsed time is computable against the SLA the control names. | Today’s user list showing the person is absent. That proves current state; it evidences neither the act of revocation nor when it happened. |
| Quarterly access review | The extract that was reviewed, the reviewer’s identity, the decisions recorded line by line, and the downstream tickets showing the removals were actioned. | A signed-off review with nothing showing anything was revoked. That evidences the review and not the control — the control is the review plus the action it produces. |
| Production change | The change record showing the approver’s identity and an approval timestamp earlier than the deployment timestamp, plus the diff or pull-request link the approval attaches to. | An approval recorded by the same person who authored the change; or an approval carrying a date but no time, on a control whose wording says “before deployment”. |
| Vendor security review | The completed assessment, dated before the contract effective date or go-live date — whichever the control names — with the reviewer and the disposition visible. | An assessment dated after go-live. A review performed once the vendor is already processing data is a different control from the one described. |
| Alert triage | The ticket showing detection time, first-touch time, the triage notes, the severity assigned and the closure rationale — enough for the auditor to see that judgement was applied. | A closed ticket with no notes. Closure is a status change; the control is the investigation, and an empty ticket evidences the status rather than the work. |
| Annual recovery test | The test plan, a dated execution record, the results including what failed, and the follow-up actions raised out of it. | A plan document with a review date and no execution evidence. A reviewed plan evidences design; the tested control is the execution. |
One rule sits underneath all of it: the control’s performance — and the artefact recording it — has to be contemporaneous. The most expensive habit in a first Type 2 is performing controls diligently and evidencing them retrospectively during fieldwork — a reconstructed access review carries a fieldwork date, and that date records when the artefact was made, not when the review happened. Where TCSA runs readiness this is what we push hardest on, alongside dry-running the population exports before the observation period opens rather than discovering their filters mid-fieldwork. See common SOC 2 pitfalls for the rest of that list.
Edge cases
Where this genuinely gets contested
The material above is mechanical. These are the cases that produce the real arguments, and none has a single correct answer that applies everywhere.
- 1A tool migrated mid-period. Changes lived in one system until 14 June and another afterwards. The population is the union of both, and the migration date itself needs evidencing. A single export from the new tool is structurally incomplete — the failure mode auditors see most often after a platform change.
- 2The control changed mid-period. A new approval step, reviewer or threshold means two controls and two sub-populations in substance. Most auditors sample from each rather than pretend one control operated throughout, and the change belongs in the system description as a significant change.
- 3Service accounts, bots and contractors. Are they in the access population? The system description decides, and if it is silent the scoping conversation happened too late. Silent exclusion is a completeness failure; documented, reasoned exclusion is a scoping decision a reader can evaluate.
- 4Deleted or reopened records. A ticket closed in May, reopened in July and closed again is one item or two depending on the control text — and an export run in October may not match one run in August. Auditors increasingly ask when the export was generated for exactly this reason.
- 5Deviations found in the final week. They cannot be remediated inside the period, because the period is closed. Remediation after period end belongs in management’s response, not in the test result: a control performed after the period closes cannot evidence that it operated during it. Evidence obtained after period end about in-period operation is entirely normal — that is what fieldwork is, and the log exported in November is ordinary evidence about an October control. Evidence of a control performed after period end is not.
- 6Controls at a subservice organization. Under the carve-out method they sit outside the population entirely and the auditor tests your monitoring of the vendor instead. Under the inclusive method they are inside it, and the vendor must produce populations to the same standard you do.
- 7Statistical claims that are not statistical. Most SOC selection is non-statistical — haphazard or judgemental — even though the larger sizes trace to attribute-sampling tables built at roughly a 90% confidence level, with a tolerable deviation rate in the region of 9 to 10% and zero expected deviations. That combination is where 25 comes from. The low-frequency rows — 2 of 12, 2 of 4 — are judgemental conventions with no comparable statistical basis at all, which is precisely why guidance treats low-frequency populations as a matter of judgement. A firm quoting a confidence level without also quoting a tolerable rate is quoting half a parameter; a firm describing a haphazard selection as “statistically valid” is overstating it. Both are fair questions to ask.
Objections
What a buyer’s security team pushes back on
“You only tested 25 of 4,000 changes. That proves nothing.”
Sampling is the normal way tests of controls are performed where the population is too large to examine in full; where it is small enough, the auditor examines all of it. The conclusion rests on the population being complete and the selection being independent, not on volume. The useful questions are the auditor’s: was the population validated, who selected the items, and what were the results. Twenty-five items from a validated 4,000 is worth more than 200 from an unvalidated 2,000.
“Your monthly control was only tested twice.”
Two of twelve is the conventional extent at an expected deviation rate of zero, set by the CPA firm’s methodology rather than by the client. The more searching question is whether the population really was twelve. A monthly control with a population of nine has already failed three times before sampling started — and that is the finding worth looking for.
“There is an exception in Section 4 — the control failed.”
An exception is a described deviation, evaluated by the auditor and disclosed. Read it with the opinion: under an unmodified opinion the auditor concluded that the controls, taken as a whole, operated effectively to provide reasonable assurance that the applicable trust services criteria were achieved, notwithstanding what they found. The opinion is not expressed criterion by criterion, so a single exception is not a per-criterion verdict. A report with no exceptions at all across a large control set and a twelve-month period is the one worth a second look.
“Why does this control show no testing at all?”
Almost always a nil population — the triggering event did not occur during the period. It is disclosed rather than hidden, and the auditor should have satisfied themselves the zero is real. It does mean you receive no assurance about that control, which is a legitimate thing to weigh and to ask the vendor to address next period.
“Your evidence came out of a compliance platform, not your auditor.”
The platform’s listing is information produced by the entity like any other export, and it is treated exactly that way. The auditor still has to establish how the platform derived it: which integrations were connected, which accounts, repositories, cloud resources and directories were in scope, when each integration was connected relative to the period, and whether anything was excluded by a filter or a suppressed check. Selection of the items remains the auditor’s. A platform shortens evidence collection; it does not transfer the completeness question to the vendor.
“Who decided what counted as a change?”
The control text and the system description decide, and both are management’s assertion — examined by the auditor, not authored by them. So the sharper question is whether the population definition matches the system boundary set out in Section 3: does any deployment path, environment, region or product line sit outside what the description says the system is? A population that is complete against a narrow boundary is still a complete population. It is the boundary, not the sampling, that limits what the report covers, and that is where a buyer’s scrutiny is best spent.
A closing note on roles, because sampling questions often get pointed at the wrong party. The methodology belongs to the licensed CPA firm that performs the examination under AT-C 105 and AT-C 205 and signs the opinion — SOC 1 works the same way under AT-C 320, with control objectives in place of the trust services criteria. Tranquility Cybersecurity does readiness, control design, evidence architecture and population preparation, and coordinates the examination; we never test, attest, certify or sign. If your auditor’s extent of testing puzzles you, asking for their methodology is entirely reasonable.
SOC 2 Sampling — Common Questions
Populations, extent of testing, deviations, and the awkward cases.
What is a population in a SOC 2 audit?
A population is the complete set of items a control was applied to during the observation period, and its shape comes from the control language. A scheduled control — a quarterly access review, a monthly reconciliation — has a population of its own occurrences: four, or twelve. An event-driven control has a population of events: every joiner, every leaver, every deployment, every alert above the threshold the control names. The auditor fixes the population before selecting from it, because a sample only supports a conclusion about the population it came from.
Why does the auditor test population completeness before sampling?
Because a sample says nothing about items that were never in the listing. If forty people were provisioned during the period and the export shows thirty-two, testing five supports a conclusion about those thirty-two only — and the eight missing records are exactly where an unapproved account is most likely to be. An incomplete population does not produce weak evidence; it produces evidence about the wrong set. It is also why a completeness failure found after testing forces every tested item to be re-evaluated.
Is there an official AICPA sample size table for SOC 2?
No. Neither the AICPA nor any other body mandates sample sizes for SOC engagements. The frequency-based ranges everyone quotes — two of twelve for a monthly control, five to nine of fifty-two for a weekly one, twenty-five for a population above two hundred and fifty — are CPA firm methodology, derived from the AICPA’s audit sampling guidance for financial statement audits and adopted into SOC practice by convention. AT-C 105 and AT-C 205, as recodified by SSAE 21, set none of them. Treat them as market practice: what justifies a firm’s numbers is a documented, consistently applied methodology, not the table.
How does a shorter observation period affect sample sizes?
Populations scale to the period, so samples scale with them. In a three-month Type 2 a monthly control has a population of three rather than twelve, and a quarterly control may have a population of one. The extent of testing falls accordingly, which means a short first Type 2 delivers materially less tested evidence than a twelve-month one — separately from the fact that a short report ages faster in front of buyers. There is a second effect: a twelve-month examination is usually tested in two instalments — interim fieldwork a few months before period end, then a roll-forward covering the remainder — whereas a three-month period rarely justifies the split, so all of its testing lands after the period has closed. That consequence belongs in the decision alongside cost and timing.
What is the difference between a deviation and an exception?
A deviation is what the auditor finds during testing: a sampled item where the control did not operate as described. An exception is what appears in Section 4 of the finished report. The two are separated by an evaluation — the auditor investigates nature and cause, tests whether the failure is confined to an identifiable sub-population, and assesses the effect on the criterion. Some deviations dissolve on investigation; those that do not are disclosed as exceptions. Where the opinion is unmodified, the auditor has still concluded that the controls, taken as a whole, operated effectively to provide reasonable assurance that the applicable trust services criteria were achieved — the opinion is not expressed criterion by criterion, so an exception is not a per-criterion verdict.
Does one deviation mean the control failed?
Not automatically, and the honest answer depends on the population. In six hundred production changes, one deviation in a sample of twenty-five is a four per cent rate, and the auditor can usefully ask whether it is confined to a particular route, person or fortnight — commonly by identifying every item sharing that feature and testing all of them. In four quarterly reviews with a sample of two, one deviation is fifty per cent and little room remains. In an annual control it is one hundred per cent.
Can the auditor extend the sample after finding a deviation?
Yes, and it is common — a typical firm pattern for a large population runs twenty-five items, extends to forty after one deviation and to sixty after a second, then stops extending, at which point most methodologies conclude the control did not operate effectively for the period. That pattern is market practice, not a rule anyone is bound by. Extension gathers information about whether the failure was confined; it never erases the original deviation, which stays in the numerator. Asking an auditor to test more items in the hope the average improves misunderstands the exercise, and substituting a failed item for a different one is not available at all — selection belongs to the auditor.
Why was one of our controls tested only once?
Almost certainly because it is a configuration rather than an activity — an enforced password policy, mandatory multi-factor authentication, a session timeout, a bucket blocking public access. The system applies the same rule to every instance, so inspecting the configuration once can evidence how it behaved across the period. That sufficiency is borrowed from the general IT controls: it holds only where change management (CC8.1) and restricted administrative access (CC6.1, CC6.3) are themselves reliable.
What happens if a control had no occurrences during the period?
For a genuinely event-driven control — no incidents, no terminations, no emergency changes — this is a nil population and is reported as such. Section 4 typically records that no instances occurred and that no operating-effectiveness testing was performed, and some auditors add a passage to the auditor’s report identifying controls that did not operate. Two cautions: the zero must be evidenced rather than asserted, because a zero from a wrong filter looks identical to a real one; and a scheduled control that did not run is a deviation, not a nil population.
Can we tell our auditor which items to test?
No, and offering to is counterproductive. Selecting the items is the auditor’s work, and the independence of that selection is part of what makes the conclusion meaningful — a sample chosen by the party being examined supports nothing. What you may reasonably ask is which method was used: random selection with a randomising device, which is the only one that supports statistical projection; systematic selection at a fixed interval after a random start; haphazard selection made without a device and without conscious bias, which is the most common on SOC engagements and is not statistical; or judgemental selection weighted toward higher-risk items. What you should do is prepare the population: dry-run the exports before the period opens, confirm the filters capture every path, reconcile counts to an independent source, and create evidence as the control operates rather than reconstructing it during fieldwork.
Related reading: how to read a SOC 2 report, opinions and exceptions, choosing your observation period, the SOC 2 controls list, scope and system boundary, and the SOC 2 hub.
Written By Expert Auditors
Keep Exploring
Related Reading
SOC 2 Evidence: What Auditors Accept
Why screenshots are weak, what a defensible artefact contains, and the common rejections.
Read moreSOC 2 Controls List
No official list exists — illustrative controls by criteria series (CC1–CC9).
Read moreSOC 2 Opinions & Exceptions
Unmodified, qualified, adverse, disclaimer — and why exceptions are not qualification.
Read moreSOC 2 Observation Period: 3, 6 or 12 Months?
No AICPA minimum. Why short windows leave low-frequency controls untested.
Read moreSOC 2 Knowledge Hub
Type 1 vs Type 2, criteria, timelines and audit prep — all guides.
Read moreSOC 2 With a Small Team
Segregation of duties when the founder approves everything — the compensating controls that actually pass.
Read moreGet in touch
Book a free consultation or send us your requirements. We respond within 24 hours.
Quick Call
Pick a time slot
Send Requirements
Get a custom quote in 24 hours