Audit testing is not improvisation - it uses four recognised methods, ranked by how much assurance each provides. Knowing which method applies to which control tells you precisely what to prepare, and explains why some requests feel disproportionate to the control being examined.
The four methods, weakest to strongest
1. Inquiry - asking
The auditor asks how something works. It is the weakest form of evidence and is never sufficient alone, because it captures belief rather than fact. Its real purpose is orientation: it tells the auditor what to go and verify.
Practical consequence: a confident verbal answer that later contradicts a document is far worse than saying "let me check". Inconsistency between what people say and what records show is the single fastest way to expand an audit's scope.
2. Observation - watching
The auditor watches the control being performed - a badge check at reception, a deployment approval, a review meeting. Stronger than inquiry, but it proves only that the control operated while being watched, which is precisely its limitation.
3. Inspection - examining records
The workhorse. The auditor examines artefacts the control produced: approval tickets, review exports, meeting minutes, configuration screenshots, logs. This is where most audit evidence comes from, and why "the control happens but leaves no record" is treated as a failure rather than a documentation nicety.
What makes an artefact inspectable: it is dated, it identifies who performed the action, it shows what was decided, and it comes from a system the performer cannot silently alter.
4. Reperformance - doing it again
The auditor independently re-executes the control and compares results - re-running an access export and checking it against your reviewed list, recalculating a figure, re-testing a restore. It is the strongest method because it does not rely on your record being complete or accurate.
Reperformance is where discrepancies surface that inspection misses: accounts absent from the export you reviewed, exceptions filtered out before the report was produced.
Have the evidence ready before it is asked for
GRC Copilot files dated, control-mapped evidence as it is produced, so sample requests are answered from a search rather than a scramble.
Try GRC Copilot free Generate an AI-powered assessment Download checklist Book a demo
Population completeness comes first
Before sampling anything, a competent auditor establishes that the population is complete. If you provide a list of 340 changes for the period, they will want comfort that 340 is all of them - typically by pulling the list themselves, reconciling to a system count, or checking that the extract had no filters applied.
This is the step organisations underestimate. A perfect sample drawn from an incomplete population is worthless, and an auditor who discovers the population was filtered will question every test that depended on it.
Practical implication: never hand over a manually curated list. Provide the raw export and let the filtering be visible, even when it includes items you would rather explain.
Sampling
Auditors sample rather than test everything, and sample size scales with how often the control runs. Common guidance looks roughly like this:
- Annual control - 1 instance.
- Quarterly - 2.
- Monthly - 2 to 5.
- Weekly - 5 to 15.
- Daily - 15 to 40.
- Many times per day / automated - 25 to 60, or a single configuration test plus evidence the configuration did not change.
Two things follow. First, an automated control can often be tested once rather than sampled dozens of times - a strong argument for automating high-frequency controls. Second, exceptions are not scaled: one failure in a sample of two is treated very differently from one in forty, and a small sample gives you almost no room for error.
What happens when a sample fails
The auditor will usually expand the sample rather than immediately raise a finding. If the expanded sample is clean, the original may be treated as isolated - documented, but not necessarily a control failure. If failures recur, the control is deemed ineffective for the period, and remediating it now does not retrospectively fix the period tested.
The useful response to a failed item is a factual explanation with supporting evidence, offered immediately - not an argument about whether it counts.
Frequently asked questions
Can we choose the sample?
No - selection must be the auditor's, or the test provides no assurance. You can and should provide the complete population promptly.
Why do they want the raw export rather than our summary?
Because a summary is your assertion, not evidence. The export lets them establish completeness independently, which everything else depends on.
Is a screenshot acceptable evidence?
Sometimes, for configuration. It is weak - undated, croppable, and easy to stage. System-generated exports and tickets are far stronger, and a screenshot showing the timestamp and the full screen is much better than a cropped one.
What if evidence for one sampled item genuinely does not exist?
Say so immediately and explain why. Reconstructing or backdating a record turns a control exception into an integrity issue, which is a categorically worse conversation.
Key takeaways
- Four methods, ascending in strength: inquiry, observation, inspection, reperformance.
- Inquiry alone never suffices - it directs the auditor to what they will verify.
- Population completeness is tested before sampling; hand over raw exports, not curated lists.
- Sample size scales with control frequency, which makes automation disproportionately cheap to test.