Auditors ask two independent questions about every control: is it capable of achieving its objective, and did it actually happen? The first is design effectiveness, the second operating effectiveness. Passing one tells you nothing about the other, and confusing them is why teams remediate the wrong thing.
Design effectiveness
Design asks a hypothetical: if this control operates exactly as described, would it prevent or detect the thing it is meant to address? The evidence is the control's definition - the policy, the procedure, the system configuration, the workflow - not its history.
Typical design failures:
- The control does not address the risk. Quarterly access reviews do not detect a privilege granted and abused within a week. The control is real; it just does not do the job claimed for it.
- The frequency is wrong for the exposure. Annual reviews on a system with weekly joiners.
- The performer cannot be objective. The person who granted access reviewing the access.
- No output. A control that produces no record cannot be tested by anyone, including you.
- Scope holes. The control covers production but not the data warehouse holding the same data.
- Undefined trigger. "Reviewed when significant changes occur" with no definition of significant.
Design failures are the expensive ones. Operating failures cost you a sample; design failures mean the control was never going to work, so every period it has run is also questionable.
Operating effectiveness
Operating asks a factual question: did the control run as designed, consistently, throughout the period? The evidence is history - dated records, tickets, exports, approvals, logs. Typical failures:
- It ran three quarters out of four.
- It ran, but the output shows exceptions nobody actioned.
- It ran only for part of the population.
- It ran late, repeatedly, against its own stated frequency.
- It ran, but produced no evidence - which auditors treat as not having run.
Prove both, continuously
GRC Copilot records how each control is designed, schedules its operation, and files the dated evidence that proves it ran - so both tests are answered from one place.
Try GRC Copilot free Generate an AI-powered assessment Download checklist Book a demo
Where you meet the distinction formally
- SOC 2 Type I tests design only, at a point in time. Type II tests design and operating effectiveness across a period - which is why Type II requires months of accumulated evidence and Type I does not.
- ISO 27001 Stage 1 largely examines design and documentation; Stage 2 examines operation.
- Internal audit conventionally performs a walkthrough to assess design, then sample testing to assess operation - and stops if design fails, because testing the operation of a broken design is wasted work.
Why the order matters
Design is assessed first, deliberately. If a control is not capable of achieving its objective, the fact that it ran twelve times is irrelevant - you have consistently performed something that does not work. Auditors will not spend sample effort on it.
This has a practical consequence for remediation. When a finding arrives, establish which test it failed:
- Design failure - change the control. New frequency, new performer, wider scope, a defined trigger, an output that can be evidenced. Then start accumulating evidence again from zero.
- Operating failure - the control is right; the discipline is not. Assign an owner, schedule it, monitor completion. Do not redesign a working control because it was skipped once.
Teams routinely do the opposite: they rewrite a sound control after a missed quarter, resetting the evidence history for no reason.
The awkward middle case
Sometimes a control operates perfectly and still fails design because it was never capable of covering the whole population - for example a review that only ever looked at one of three identity providers. The records look complete because the population was defined too narrowly. This is why auditors test population completeness before sampling: an accurate sample of the wrong population proves nothing.
Frequently asked questions
Can a control pass design and fail operating?
Very commonly - it is the single most frequent finding pattern. The control is sound and somebody stopped doing it.
Can it pass operating and fail design?
Yes, and it is worse. Consistent performance of a control that cannot achieve its objective means the risk was never addressed in any period.
How do we self-test design?
Take the risk, read the control aloud, and ask what a determined insider or a plausible mistake would do. If the control would not catch it, the design is weak regardless of how well it is documented.
Does automation solve operating effectiveness?
Largely, yes - an automated control that runs on a schedule rarely gets skipped, and it logs itself. It does nothing for design; an automated bad control is a reliably bad control.
Key takeaways
- Design asks "would it work?"; operating asks "did it happen?" - independent tests.
- Design is assessed first; a failed design makes operating evidence irrelevant.
- Fix design failures by changing the control, operating failures by fixing ownership and cadence.
- Check population completeness - a clean sample of the wrong population proves nothing.