You’re looking for the deviator who failed. The deviator was doing what the system trained them to do.


Bainbridge’s irony from chapter 13, that automation hollows out the operators it relies on, runs into incident response next, where blaming the operator leaves the conditions that produced the incident untouched.

🎧 Prefer to listen? This chapter is narrated in my own voice with ElevenLabs on Spotify (30 minutes).


The two poles

Pole A: deviation. When work goes wrong, the explanation is a deviation from procedure, competence, or intent. The fix clarifies the standard or addresses the deviator. Human error is a sufficient diagnosis.

Pole B: conditions. Real work continually adjusts to underspecified conditions. The gap between work-as-imagined and work-as-done is where both safety and risk live. Failure is the unexpected combination of normal variability. Human error is a symptom of that gap, not a diagnosis.

Table comparing Deviation (Dekker's Old View) with Conditions (Dekker's New View) across four rows: human error, what failed, the fix, and the question asked.
The two poles side by side: what each account treats as the diagnosis, what it says failed, what it fixes, and the question it asks in the review.

Where Pole A is right

Pole A fits tractable failures with a clean mechanism and locatable responsibility: a broken bearing, a software null-pointer exception, a structural calculation error, or an individual whose behaviour was genuinely reckless. Even that last item carries Dekker’s caveat. A proven bad apple is still a system responsibility: someone hired, credentialed, and scheduled them. Further, reckless is an infinitely negotiable label through which the deviation reflex re-enters, so treat it as where an investigation starts, not as its conclusion.

Pole A is right when the failure is a discrete event in an otherwise stable system and responsibility can be named without distortion. Root-cause analysis works for that kind of failure. Many engineering-organisation failures are not that kind.

Where Pole B is right

Pole B fits any incident where failure came from the unexpected combination of normal variability, which, in my experience, is most of what goes wrong in the complex sociotechnical systems senior engineering leaders actually run. The cascade where every service degraded inside its own tolerances and the aggregate still took the platform down. The data-loss incident where the backup job and the restore path were each correct and had never been run against each other. The retry storm where the autoscaler, the client timeout, and the retry policy all behaved exactly to spec.

In decisions

Pole A leaders trace the incident back to the decision that broke the standard, then reinforce the standard and coach the person who missed it. Pole B leaders go to the gemba—the actual place where the work happens—and ask “what surprised you?” and “where did you have to improvise?” Then, they change the system people are operating within.

Take your last meaningful incident and lay two analyses side by side. The first is your standard after-action review. The second report is what the on-call rotation knew beforehand and had no reliable channel to surface. That second report doesn’t exist until someone goes to the gemba and sources the knowledge, which takes effort and trust. The full protocol is at the end of the chapter.

The gap between the two analyses is what Vaughan calls structural secrecy, a mechanism this chapter comes back to. Ask what this organisation has collectively stopped being able to see.

The after-action review sits inside your own span, and that part doesn’t have to be an ask: I intend to redesign our reviews so the broader context is examined every time. Our current review isn’t going deeply enough, so unless you see something I don’t, the redesigned review starts with the next incident. The sentence your CEO can carry to the board: “What reaches us has been filtered; the deeper, unexamined context is where the next incident is hiding.”

The Old View and the New View

Sidney Dekker, in The Field Guide to Understanding ‘Human Error’, names two epistemic frames. The Old View holds that human error causes accidents. People who err are bad apples: remove them, retrain them, or stiffen the rules. The New View holds that human error follows from trouble inside the system. People who err are the symptom; the conditions that shaped their behaviour are what needs explaining.

Dekker’s case, made across 30 years of safety-science work, is that the Old View’s empirical props—accident-proneness statistics and Heinrich’s triangle—fail under the data, while its remedies suppress the information safety management needs.

Heinrich’s triangle holds that minor incidents and fatalities sit in fixed proportion, so a falling count of small events implies the big ones are under control. Incident dashboards run on that logic. Heinrich built the triangle from insurance and actuarial science: the ancestor of the modern incident dashboard is an underwriter’s loss table from nearly a hundred years ago. Dekker’s counter-example is Deepwater Horizon, where managers were celebrating six years of injury-free performance the day before the rig killed 11 people and caused the largest accidental marine oil spill in history. The record was real and accurately measured whether people were getting hurt on deck, but the well was a different risk class entirely. Construction shows the same split over decades: minor-injury frequency falling while the absolute count of fatalities and life-altering injuries holds steady. The ratios vary wildly between industries, jobs and places, and the safer a system gets, the less the small stuff predicts the large. Dekker’s charge is stronger than inaccuracy. The belief sedates: it “rocks you to sleep with the lullaby that the risk of major accidents or fatalities is under control as long as you don’t show minor injuries, events or incidents.” The quotation marks in Dekker’s book title compress the argument into typography: human error is a label someone applies, and labels are not categories of behaviour.

The other prop is older, and it keeps coming back. In the mid-1920s the Boston public transport company found that 27% of its drivers caused 55% of its accidents, and psychologists in Britain and Germany independently proposed that some workers were simply accident-prone. Every engineering leader has seen the software version: one name recurs across the incident log, and the pattern looks like a person. The refutation is a denominator problem. The method needs every driver to face the same expected accident rate, and some drove the busy centre while others drove quiet suburbs at night. Dekker states it flatly: practitioners are not all exposed to the same kind and level of accident risk, which makes it impossible to compare their rates and conclude that personal characteristics explain the gap (p. 12). The engineer who appears in every incident review usually owns the gnarliest service. By the end of WWII the field had concluded that proneness was “much more a function of the tools and tasks he or she was given, and the situations they were put into” than of personal characteristics. The idea survives because blaming the worker leaves management “scot-free” of responsibility for design, selection and training.

The Old View persists because it gives leaders something to do that feels like action and puts the accountability outside the decisions and conditions that are theirs alone. Naming the deviator is easy and fast. Understanding, naming, and resolving the conditions is hard and slow.

Naming the deviator is easy and fast. Understanding, naming, and resolving the conditions is hard and slow.

Practical drift, at four levels

Scott Snook’s Friendly Fire documented the 1994 friendly-fire incident in northern Iraq. Two American F-15 pilots, operating under Operation Provide Comfort, mistakenly engaged and destroyed two American Black Hawk helicopters. 26 people died. I think Snook’s book is the single most important book-length review of how organisational accidents actually happen.

Snook refuses to pick one level like a Pole-A reading would. The pilots fired the missiles, so the pilots are the cause. Or the AWACS crew failed to challenge the engagement, so the AWACS crew is the cause. Or the rules of engagement were ambiguous, so the rule-makers are the cause. Each was true at its level, but none is the whole story.

Snook’s four-level causal map names the structure: individual (what the pilots saw and did), group (what the AWACS crew did), organisational (how the no-fly zone command structured decision rights), and cross-level (how the layers interacted under conditions of practical drift). Each layer has its own causal story, none sufficient on its own to explain the accident.

Picking one level is, in Snook’s framing, how resolution gets felt without being reached. The simple story picks one level and calls it the cause; the multi-causal story holds all four at once.

Snook named the pattern of small-step accommodation across these levels practical drift: “the slow steady uncoupling of local practice from written procedure”. The 2×2 he drew has two axes: is the action rule-based or task-based, and is the system loosely or tightly coupled. Four quadrants follow: Designed, Engineered, Applied, Failed. Most organisations live in the one Snook called Applied, where local task logic and loose coupling keep everything working.

Two-by-two of rule-based versus task-based action against loose versus tight coupling, giving Designed, Engineered, Applied and Failed, with Applied highlighted as where most organisations live.
Snook’s practical-drift matrix: rule-based against task-based action, loosely against tightly coupled. Most organisations sit in Applied, one coupling-tightening away from Failed.

Normal Behavior, Abnormal Outcome is his heading for what the move from Applied to Failed produces when coupling suddenly tightens. Everyone is doing what local practice says. The aggregate produces an outcome no one actively chose. The matrix is a cycle: the organisation that has just been to Failed writes tighter rules, which is the move back to Designed, and the drift begins again from there.

Snook’s verdict is that the shootdown was, in the strict sense, normal. Normal people behaved in normal ways inside normal organisations, exactly as the theory would predict given the circumstances each level faced.

The Pole A response to practical drift looks for the moment someone broke a rule. There was no single decisive moment. Pole B asks one pair of questions at all four levels: what did work-as-imagined say should happen, and what had work-as-done drifted into?

Structural secrecy: what makes drift invisible

Snook’s practical drift named the pattern. Diane Vaughan’s The Challenger Launch Decision names what makes the pattern hard to recognise as a signal from inside the organisation.

Structural secrecy is Vaughan’s name for the way an organisation’s own patterns of information, structure, processes and regulatory relations undercut its attempt to know and interpret its own situation. That runs at every level, not only the top. It is the ordinary functioning of a multi-level organisation, without concealment by intent.

Vaughan’s three social forces from the Challenger analysis—production of culture, culture of production, and structural secrecy—together explain why deviance gets normalised. Structural secrecy is the central one for a senior leader’s purposes. In Vaughan’s own case the channels were open and the top knew what the work group knew; the filter was as much in the shared construction of risk as in transmission.

Dekker carries Vaughan’s concept into the safety bureaucracy itself: the cultural, organisational, physical and psychological separation between operations on one side and safety regulators, departments and bureaucracies on the other, a gap widened by bureaucratic entrepreneurism. Clarke and Perrow’s 1996 fantasy documents sit in that gap: safety plans that bear no relation to actual work, written to persuade regulators and boards that the hazard is handled. They are the paperwork that makes an organisation feel like it can see.

Structural secrecy and practical drift feed on each other. Not knowing what other units do lets work groups drift into locally practical arrangements, and the distance that opens between groups deepens the secrecy in turn.

Cook’s 18 points

I suspect Richard Cook’s How Complex Systems Fail (a short treatise first published in 1998) is the densest one-document statement of the New View. It contains 18 numbered claims. Three matter most for this axis; the point titles below are my paraphrases, not Cook’s.

Point 7: There is no isolated root cause for failure in a complex system. Failure emerges from multiple latent conditions interacting under specific circumstances. The search for a single root cause is, in this framework, a categorical mistake. The system produces failure the way a wet floor produces a slip. The floor alone is harmless. It takes the wet floor, the smooth shoes, the box someone is carrying that they can’t see past, and the hurry, all present at once, and the slip is what that combination does.

Point 8: Hindsight biases post-accident assessments. After the fact, the path to failure looks obvious. Why didn’t they see it? They couldn’t see it from inside the situation because the path became obvious only in retrospect. Further, I’ve observed that the longer the gap between incident and review, the simpler and vaguer the story gets and the less the story tells you about what to actually change.

Point 15: Post-accident remedies often increase coupling and complexity. The intuitive Pole A response—add a procedure, add an approval gate, add a monitoring system—frequently makes the system more tightly coupled and harder to operate safely. The remedy can make the next incident more likely. Cook’s claim is hard to act on because the alternative—sit with the incident, study the conditions, do less—violates Pole A’s instinct. “Don’t just do something, stand there (in the Ohno circle)!”

Point 15 names a design choice as well. You can try to make the system unable to fail by listing every way it could fail and preventing each one, or you can design it to contain failure when it arrives, because no list is complete in a self-organising system. Alicia Juarrero’s constraint architecture, developed at length in the local-optima chapter, is the philosophical statement of the same choice. Punishing the operator leaves that architecture untouched, which leaves the next incident pre-loaded.

Behind Human Error

Most incidents carry two stories. The first story is the leader’s cover for the system they are responsible for: the operator failed to follow procedure. The second story is the one underneath it. The procedure didn’t fit the conditions the operator was facing; the operator improvised; the improvisation worked most of the time… until it didn’t.

Keep working until you have the second story. Sharp end—the operators in the moment—and blunt end—the organisational conditions shaping what the sharp end can do—must both be in the analysis. Punishing the sharp end while leaving the blunt end unchanged is how after-action review most often fails.

Punishing the sharp end while leaving the blunt end unchanged is how after-action review most often fails.

The boundary the organisation draws between operator and system is political, and that choice determines who gets blamed. Pole A draws the line so the operator stands alone. Pole B includes the system in the analysis.

Work-as-imagined and work-as-done

Above the line is what the organisation describes about how the work gets done: org charts, policies, procedures, dashboards. Below the line is what actually happens when the work is being done: the improvisations, the workarounds, the implicit knowledge, the social adjustments. The gap between the two is where safety lives, and where risk lives. (The phrase is adapted from the STELLA Report’s line of representation, which draws its line in a different place: the people and their mental models above it, the code and hardware below.)

Diagram with org chart, policies, procedures and dashboards above a horizontal line, and improvisations, workarounds, implicit knowledge and social adjustments below it, with the gap marked.
Above the line is what the organisation describes. Below it is what people actually do. The gap is where safety and risk both live.

The pair of terms is Erik Hollnagel’s. After-action reviews that stay with work-as-imagined produce policy clarifications. Reviews that reach work-as-done produce systemic learning. Spend enough time in the work for work-as-done to become legible.

Vaughan’s structural secrecy operates here. The organisation’s above-the-line representation is the documented work. The below-the-line reality is the actual work. Structural secrecy lets both versions coexist without anyone recognising the gap as a signal.

Reason’s informed culture

James Reason names the cultural toolkit Pole B requires. An informed culture—one in which leadership knows what is actually happening at the sharp end—rests on four interlocking subcultures: a reporting culture where people freely surface errors and near-misses without fear of inappropriate blame; a just culture that distinguishes honest error from the small minority of behaviours that warrant sanction so the reporting culture can survive; a flexible culture that can move authority to the sharp end during high-tempo operations; and a learning culture that turns lessons into reconfigured assumptions and action.

Reason’s just culture carries a caveat. It assumes the line between honest error and sanctionable behaviour can be drawn clearly and consistently. Dekker’s Just Culture argues there is no line, only people with the power to draw it: the same negotiability that hangs over reckless earlier in this chapter. Treat the distinction as something your organisation negotiates in the open. The taxonomy will not hand it to you.

Engineering organisations routinely claim what Ron Westrum called a generative culture and operate, in fact, as bureaucratic ones. Reason’s framework tells leaders which subcultures must be present before that claim is empirically true.

Reason’s separate GEMS taxonomy supports a more precise discipline. The same outcome—the operator did the wrong thing—has at least three different mechanisms underneath, and each needs a different response. A skill-based slip or lapse, where the action didn’t match the intention, wants better interface design and reduced attentional load. A rule-based mistake, where the wrong rule fired, wants better rules and better training in when each applies. A knowledge-based mistake, where no rule applied and the operator had to improvise a plan that turned out wrong, wants better support for genuine sense-making in novel territory.

Diagram splitting 'the operator did the wrong thing' into a slip, a rule-based mistake and a knowledge-based mistake, each with its own response: redesign the interface, better rules trained to their boundaries, and support for sense-making in novel territory.
One outcome, three mechanisms. A slip, a rule-based mistake and a knowledge-based mistake each want a different response. The deviation reflex applies the same remediation to all three.

The Swiss cheese model is itself a simple story

Deviation and conditions are also two ways of telling the story of a failure. The deviation account is a simple story: one cause, one villain, one fix. The conditions account is multi-causal: several causes operating at once, no single villain, a fix that touches the system. The narrative form determines the intervention that follows, which is Jennifer Garvey Berger’s simple-stories mindtrap operating on an incident review.

James Reason gave the field what became known as the Swiss cheese model: defensive layers with holes that vary over time, and accidents happening when the holes align. The metaphor crossed every tradition in the field and is the most travelled single framing of complex-system failure in the literature. Unsafe acts by front-line workers still do real work inside Reason’s model, and the phrase is Heinrich’s. Dekker dates the thinking underneath the diagram as “getting on in age—soon a century” (p. 123).

Four slices labelled design, training, procedure and the operator, each carrying holes, with a single arrow running left to right through the aligned holes.
The Swiss cheese model draws multi-causal failure as a straight line through four slices. Use it as the entry point, then ask what a straight line cannot draw.

The metaphor itself tells a simple story about multi-causal failure. Linear slices. Discrete holes. Causation moves left-to-right through the diagram. The model meant to teach multi-causality is the most Pole-A-shaped way to talk about it.

Even the field that argues hardest for multi-causal stories reaches for a linear metaphor when it has to communicate. The Swiss cheese model travels because it is simple. Its simplicity makes it portable and limits what it can teach.

The Pole B move is to use the model as an entry point and refuse to stop there. When the after-action review reaches for the Swiss cheese diagram, ask which non-linear interactions, dynamic feedback loops, and structural conditions the linear diagram couldn’t depict. Carry Garvey Berger’s own summary into that room: “in a complex world a simple story is just about always wrong, and will just about always lead us to an emaciated, impoverished set of choices” (p. 39). Her habit is the discipline that room needs: notice your story, then create another, then another.

A representative case

Picture a platform engineering team a few months into a run of deployment outages. The Pole A response from leadership is clear: tighter change-control, more approval gates, after-action reviews focused on what each on-call engineer did wrong. The outages continue. The engineers are exhausted and starting to leave. The ones who were on call when it broke carry it hardest: Dekker calls practitioners harmed by an event they were caught up in second victims, and support for them is an obligation.

When the Pole B move lands, a senior engineer is given a week to do nothing but interview the on-call rotation and the platform team about what the work actually looks like below the line. What the interviews surface is consistent. The change-control system has become so onerous that engineers are batching small changes into larger ones: fewer change reviews, but each change touches more surface. The approval gates have introduced a delay that pushes deployments into peak-traffic windows. The after-action reviews have become defensive theatre that makes engineers reluctant to surface near-misses. Reason’s reporting subculture, which the leadership team thought it had, is structurally absent below the line.

Every one of the Pole A remedies is making the system more fragile. That is Cook’s Point 15 in operation. The Pole B response—reduce change-control friction, redesign after-action review to surface conditions rather than assign blame, give the on-call rotation slack, build a real reporting subculture by making the engineers who surface near-misses visibly protected—is the one with a claim on the outage rate. In this case, nothing about the engineers needs to change. The system changes. And blame-free still carries accountability: in Dekker’s terms, the account stops being something the operator settles and becomes something the operator tells. Engineers in a blameless review are, in John Allspaw’s phrase, very much on the hook for helping the organisation become safer.

The diagnostic move

Three questions for last Tuesday’s incident:

  • Which pole was I claiming? Did I treat the incident as a human deviation or systemic condition?
  • Which pole would my actual response show? Did I ask who failed? or What surprised you? Where did you have to improvise?
  • Which pole does this incident actually require? If the failure was the unexpected combination of normal variability, Pole A is structurally inadequate. If it was a single clean cause with a locatable mechanism, Pole A is apt.

The exercise

Run a second-story protocol on the team’s last meaningful incident, with one constraint: spend at least half a day below the line at the gemba before the after-action review gets written up. Talk to the operators. Ask the two questions Pole B leans on: what surprised you? and where did you have to improvise?

Map what you find at four levels—individual, group, organisational, cross-level—without forcing a single root cause. Compare it with whatever first story the organisation had already produced. The exercise lives in noticing the gap.

A second variant for organisations with mature reviews is to place the incident on Snook’s matrix. Was the action rule-based or task-based, and was the system loosely or tightly coupled? That lands you in Designed, Engineered, Applied, or Failed. Teams usually find they were in Applied and had not seen it. Naming the quadrant tells you which way the drift is running, and whether the remedy you are about to write starts the next lap.

Going upstream

In-text: the primary sources named in the body. For the Old View / New View frame, and for the dismantling of Heinrich’s triangle and the accident-proneness data: Sidney Dekker, The Field Guide to Understanding ‘Human Error’. For the multi-level anatomy and practical drift: Scott Snook, Friendly Fire. For the upstream mechanism, structural secrecy: Diane Vaughan, The Challenger Launch Decision.

Also touched: For the 18-point statement of the New View, including the no-root-cause and increased-coupling points: Richard Cook, How Complex Systems Fail. For the cultural toolkit—reporting, just, flexible, learning subcultures—and for the Swiss cheese model itself: James Reason, Managing the Risks of Organizational Accidents. On blameless review and the engineer as the person on the hook: John Allspaw, Blameless PostMortems and a Just Culture. On the negotiability of the just-culture line: Sidney Dekker, Just Culture. For the simple-stories habit of carrying two more stories: Jennifer Garvey Berger, Unlocking Leadership Mindtraps.

Go deeper: the convergence pile-ups and further reading, with no in-body anchor. The no-single-root-cause claim arrives from several traditions beyond Cook and Snook—Juarrero, Woods et al., and Hollnagel—each naming it differently. For the second story, sharp-end / blunt-end, Westrum’s pathological / bureaucratic / generative typology of organisational cultures, and interface as line of demarcation for blame (after Kelly-Bootle, 1995): Woods, Dekker, Cook, Johannesen, Sarter, Behind Human Error, second edition. For the line of representation this chapter adapts: Cook, Allspaw, Woods et al., STELLA Report. For the GEMS / skill-rule-knowledge error taxonomy: James Reason, Human Error, and Jens Rasmussen’s SRK framework, carried in Behind Human Error. For the philosophical foundation, enabling constraints and the fail-safe / safe-fail choice: Alicia Juarrero, Dynamics in Action, developed at length in the local-optima and leader-leader chapters. For negative capability, the capacity to remain in uncertainty without an irritable reaching after fact and reason: John Keats’s December 1817 letter to his brothers. Watch. Sidney Dekker, Safety Differently (2018 lecture, 33 min); software-native companion, John Allspaw, How Your Systems Keep Running Day After Day (DevOps Enterprise Summit 2017, 33 min). Read. Sidney Dekker, Drift Into Failure; Erik Hollnagel, Safety-I and Safety-II; Todd Conklin, Pre-Accident Investigations; James C. Scott, Seeing Like a State; W. Edwards Deming, The New Economics, on variation.


I work with engineering leaders on exactly this kind of paradigm work, the deeper the better. If it’s live for you, I’m happy to talk: schedule a 30-minute virtual coffee at hi.chrisgagne.com.

Some book links here are Amazon affiliate links; if you buy through them I may earn a small commission, at no cost to you.


That was the incident-level call. The system-level call asks what the organisation counts as safety. The incident counter your team reports up reads zero, and you take that for safety, when all it records is the absence of whatever got counted. It stays silent on how the work goes right on the ordinary days, which is what keeps you safe. Chapter 15 sets safety as absence against safety as presence, and reaches for what a leader would measure in its place.