
Open any struggling CAPA log and the same two phrases appear over and over: “root cause: operator error” and “corrective action: employee retrained.” Sometimes they arrive together, a matched set. And a few months later, so does the same nonconformance, because nothing about the process that produced it has changed. The operator was never the root cause. The operator was where the process ran out of defenses.
Clause 10.2.1(b) requires you to evaluate the need for action to eliminate the causes of a nonconformity so it doesn’t recur. Everything else in corrective action, the containment, the action plan, the effectiveness check, sits downstream of getting the cause right. Get it wrong and the rest of the record is well-documented wasted effort.
Why “operator error” keeps winning
It isn’t laziness, mostly. “Operator error” wins because it’s fast, it’s technically true (someone did do the wrong thing), and it closes the record. The investigation stops at the first human it finds, because going further means asking questions about training design, work instructions, fixturing, scheduling pressure, and management decisions, and those questions have owners who outrank the investigator.
A useful discipline: treat any root cause that names a person as a signpost, not a destination. If the answer is “the operator skipped the torque check,” the real question is the next one: what made skipping it possible, likely, or invisible? Was the check missing from the traveler? Is the traveler a photocopy of a superseded revision? Does the takt time assume the check takes ten seconds when it takes ninety? Can the station even detect a skipped check, or does the error only surface at final test?
Each of those answers points at a process, and processes can be fixed. People can only be blamed.
Five whys, used honestly
The five-whys method has a bad reputation it mostly doesn’t deserve. It fails in two predictable ways, and both are usage problems.
Stopping early. The chain runs “part was out of spec, because the operator misread the drawing, because they weren’t paying attention,” and closes. That’s two whys and an accusation. Keep pulling: why was misreading possible? The drawing crams four tolerance callouts into one corner at a scale nobody can read on the shop floor. Now there’s something to fix.
Branching ignored. Real failures rarely have one chain. The out-of-spec part might trace to the drawing and to the fact that receiving inspection stopped sampling that dimension in 2024. Pick one branch and the other one reoffends. Write down the branches you’re not pursuing and why; an auditor reading “second contributing cause accepted as risk, rationale attached” sees judgment. An auditor reconstructing the ignored branch from the recurrence sees a system that investigates until the answer is cheap.
The test for whether you’ve hit an actual root cause is old and still the best one available: if you removed this cause, could the nonconformance still occur by the same mechanism? If yes, keep digging. And the companion test: is the cause something management controls? “Operator error” fails that test. “Work instruction revision process doesn’t reach laminated station copies” passes it.
Proportionality cuts both ways
None of this means every NCR deserves a full investigation. Clause 10.2 says corrective action should be appropriate to the effects of the nonconformity, and a team that five-whys every cosmetic scratch will be too exhausted to investigate the field failure properly. The skill is triage with written criteria: severity, recurrence, customer impact. A one-off scratch gets a disposition and a tally mark. The third scratch from the same station gets the investigation, because recurrence just voted.
That tally mark is doing real work, and it’s the piece spreadsheet-based systems lose. Recurrence detection requires that this month’s NCR can see last quarter’s, which requires them to live in the same system with the same fields, categorized the same way. Three “one-off” events in three separate files are invisible as a pattern. The same three events as linked records against the same part and process flag themselves.
Closing the loop: effectiveness is part of the investigation
Clause 10.2.1(d) requires reviewing the effectiveness of corrective action taken, and this is where the quality of the root cause gets audited by reality. A real root cause produces a checkable prediction: “if the fixture change worked, NCRs coded to this failure mode drop to zero within ninety days.” A fake root cause produces a fake check: “verified employee was retrained. Closed.”
Schedule the effectiveness check when the action is taken, against a measurable condition, with a named owner. If the condition fails, the CAPA reopens, and the reopening is evidence the system works, not evidence it failed. The CAPAs that should worry you are the ones that closed clean and whose failure mode quietly reappeared under a different NCR category, which is one more reason the NCR log and the CAPA record need to be linked views of the same data rather than cousins who exchange holiday cards.
The one-question audit
Before your next audit, pull your last ten closed CAPAs and ask one question of each: has the nonconformance recurred? If the answer is yes for more than a couple, your corrective action process is a documentation exercise, and the root cause step is where it’s failing. The fix isn’t a better form. It’s the discipline of refusing to stop the investigation at the first person it finds, and a record system that makes recurrence loud instead of silent.