The incident counter your team reports up is at zero. What it can’t show you is how the work actually gets done safely and how to keep it that way.
Chapter 14 walked through the anatomy of incidents. This chapter turns to what gets counted, because the safety metric you use decides what gets funded.
🎧 Prefer to listen? This chapter is narrated in my own voice with ElevenLabs on Spotify (23 minutes).
The deviation-versus-conditions question asks about a specific incident. Incident counts raise a system-level question: which view of safety drives the metrics and the investment? Both matter, and each demands its own focus.
The two poles
Pole A: safety as absence. Safety is the absence of events. The fewer accidents, incidents, and near-misses, the safer the system. Safety investment is insurance against accidents.
Pole B: safety as presence. Safety is the presence of defences, the ability to succeed under varying conditions. Safety investment is investment in adaptive capacity.
Where Pole A is right
Pole A fits tractable, complicated systems with stable specifications: the lost-time injury rate on a mature manufacturing line, the field-failure count for a long-standing physical product, or the audit-finding count on a stable regulated process. The absence metric tracks what’s happening because the variability is genuinely low.
In those settings, Pole A is a strong metric because work-as-imagined and work-as-done are tightly aligned. The operations leader who uses it on a mature, stable manufacturing line is doing the right work. The same metric becomes structurally misleading in another domain.
Where Pole B is right
Pole B fits intractable, complex sociotechnical systems where work-as-done routinely adjusts to conditions work-as-imagined didn’t anticipate. A software platform under continuous deployment fits. So does a clinical service operating with staffing variability, or any system where the absence-of-events metric stops moving while everyone reports getting away with more than they used to.
In decisions
Pole A leaders count injuries. Pole B leaders ask, “How are we able to do this work well 999,999 times out of a million, and what do we need to preserve about that?”
Safety-as-presence is a capital-allocation argument wearing a safety-science coat, and that is the version a CTO or VPE should carry up to the CEO. Put one picture in front of them: the four resilience potentials—Respond (handle what’s arriving), Monitor (see what’s developing), Learn (extract the lessons), Anticipate (imagine what’s coming)—mapped against where the engineering budget actually went last year. Hollnagel is careful not to prescribe a standard balance among the four; the right mix depends on the domain. My own reading of most engineering organisations is that they are over-invested in Respond and under-invested in Anticipate (scenario planning, red-teaming, deliberate failure injection).

Few organisations choose that split; they arrive at it by inertia. Respond is urgent and visible; Anticipate is neither.
The CEO-level move is to make Anticipate a named budget line, sized as a share of what you already spend on incident response rather than asked for as new money, so it survives the next planning cycle without requiring an incident to justify it. The exercises themselves don’t wait for the budget line; running them inside engineering is your span, and that part isn’t an ask: I intend to run a failure-injection exercise each quarter, because a green dashboard tells us nothing about the failures we haven’t met. The first is a tabletop, the second runs in staging, and we go near production only once we have shown we can stop one safely. Unless you see something I don’t, the first one runs this month. The sentence your CEO can carry to the board: “Zero incidents means our people caught everything before it counted; we’re now funding the catching, not just the counting.”
Safety-I and Safety-II
Erik Hollnagel’s Safety-I and Safety-II names the framework. Safety-I—Pole A—studies what goes wrong. Its unit of analysis is the incident, and its improvement loop finds what caused the incident and seeks to prevent it from recurring. Its methods include fault-tree analysis, root-cause analysis, and near-miss reporting.

Safety-II—Pole B—studies what goes right. Its unit of analysis is the everyday successful operation. Its improvement loop seeks to understand how things work when they work, and preserve the conditions that produce success. Its methods include appreciative inquiry, work-as-done observation, and resilience engineering.
Hollnagel’s empirical claim is that Safety-I produces diminishing returns in complex sociotechnical systems. Most of the time, the system runs successfully despite gaps in work-as-imagined because people inside it make invisible adjustments. Studying only the failures misses what’s actually keeping the system safe. Safety-II keeps Safety-I and adds the missing half.
Studying only the failures misses what’s actually keeping the system safe.
The causality credo and the hypothesis of different causes
Hollnagel names the unstated assumption under Safety-I: the causality credo. In Hollnagel’s words, it is the faith that “since all adverse outcomes have causes, and since all causes can be found and dealt with, it follows that all accidents can be prevented.” Closely related is what he calls the hypothesis of different causes: the belief that things going right and things going wrong have different causes.
Resilience engineering rejects both. “Failures were the flip side of successes… things that go right and things that go wrong happen in basically the same way.”
If successes and failures share the same underlying causal architecture, studying failures alone can’t improve the system. Successes also carry useful information. They are where the adaptive capacity lives.
The Pole-B metric framework coheres only after leaders reject the causality credo. Pole B measures what works while it is working.
The four resilience potentials
Hollnagel proposes four potentials as necessary and, within the book’s argument, sufficient for resilient performance.
Respond: knowing what to do: drawing on prepared actions, adapting how the system is currently running, or improvising new responses to whatever changes, disturbances, or opportunities arrive, whether routine or not. The capacity to handle what is happening now.
Monitor: knowing what to look for: keeping watch on the things that do or could affect performance in the near term, both inside the system and out in the operating environment. The capacity to see what is developing.
Learn: knowing what has happened: single-loop learning that draws lessons from specific experiences, and double-loop learning that revises goals or objectives. The capacity to extract lessons from both success and failure.
Anticipate: knowing what to expect: looking ahead to possible disruptions, novel demands, fresh opportunities, or shifts in operating conditions further out in time. The capacity to imagine what is developing beyond what current monitoring can see.
Together, the four form what Hollnagel calls the Resilience Assessment Grid. Organisations tailor its diagnostic and formative questions to their own work, apply them repeatedly, rate the answers on a Likert-type scale, and use a radar chart to track and manage change over time. This gives Pole B an operating structure. Safety is the presence of adaptive capacity; the grid names four potentials an organisation can assess and fund.
Pole B leaders need to know which of the four potentials this organisation is strong in, which it is weak in, and how that investment changes over time. Most engineering organisations fund Respond and Learn—Respond because incidents demand it, Learn because retrospectives ritualise it—while Monitor is patchy and Anticipate is starved by default.
Cook’s points 16, 17, 18
Richard Cook’s How Complex Systems Fail gives the three central points for this axis.
Point 16: safety is something the whole system produces, not a trait of any one part. You can’t add up safe individual workers, safe individual procedures, safe individual tools, and arrive at a safe system. The safety emerges from how the parts interact under load.
Point 17: People continuously create safety. Most of what keeps the system running is invisible adjustments by people inside it. The procedure said one thing; the operator did something subtly different that worked better in this specific situation; the system survived. This happens continuously. Pole A metrics don’t see it because the metric asks did an incident happen?, and no, the incident didn’t happen, because the operator adjusted. Pole B asks the adjustment to become visible.
The procedure said one thing; the operator did something subtly different that worked better in this specific situation; the system survived.
Point 18: keeping operations failure-free depends on people who have firsthand experience of how things fail. The skills the operator uses to keep the system running come from familiarity with how it can break. Reduce failure to zero and you also reduce the operator’s familiarity with the failure modes. The system becomes brittle in the long run because the people running it lose the experience they would need to recover when something does go wrong.
The Pole A leader who succeeds completely produces a fragile organisation.
Cook’s three points and Hollnagel’s four potentials teach the same thing at two scales. Cook names what safety is in a complex system; Hollnagel names how an organisation invests in it. People continuously create safety (Cook Point 17), which Respond and Monitor make visible as moment-to-moment behaviour. Anticipate and Learn protect safety as an emergent property (Point 16).
Snook’s F-111 near miss
Scott Snook’s Friendly Fire documents an instructive counter-case to the 1994 friendly-fire incident the previous chapter described. In September 1992, a flight of Air Force F-111s nearly engaged two Black Hawks from the same Eagle Flight detachment at Bashur, identifying them only on a second pass, in the Black Hawk pilot’s testimony: “At the last minute, the guy said those look like two Black Hawks.” Two years before the eventual shootdown, the same conditions nearly produced the same outcome.

The 1992 near miss didn’t result in deaths. It also didn’t result in changes to the system. It surfaced only by chance, days later, over drinks at the Incirlik officers’ club, where the F-111 pilot told the Army aviators how lucky they were: his flight had been on the trigger and only then realised the targets were friendly. It was never reported up the chain. Nothing was done; nothing was learned.
The near miss had the same structure as the eventual shootdown: no fighter-helicopter direct comms, fighters unaware of the helicopter mission, AWACS failed to relay. Two years later, the same conditions produced the accident.
A Pole B organisation would have treated the near miss as critical data. A Pole A organisation, focussed on absence metrics, didn’t, because no accident had occurred.
The near miss was the leading indicator. The accident report was the lagging one.
Snook’s own reading is harsher still: even had the near miss been reported, he doubts it would have changed anything, because the organisational response to a report is to add rules, not to address the structural disconnects a report exposes. Reporting a near miss into a system that answers with procedure buries it a second time. Pole B demands investment in the cultural conditions where near misses get heard and the response is structural rather than procedural.
Snook’s F-111 case shows what happens without a reporting culture and a just culture, the conditions under which people can surface what they saw without fear. Engineers don’t volunteer near misses to organisations that punish their existence.
The AI overlay
AI tools sharpen a Pole-A trap that was always there, and they do so in the way chapter 13 described through Bainbridge’s ironies. The automated dashboard reads incident counts as a primary signal. It can’t see the incidents operators prevented.
As AI takes over more of the visible adjustment work, it strips out the slack the absence metric was living on. The metric never surfaced the invisible adjustments and human improvisations that kept the system running; now fewer operators remain to make them. Recovery gets harder.
Organisations risk arriving at a state where the dashboards show everything is fine and the operators report being unable to recover when anything actually breaks. AI exposes how thoroughly Pole A had already hollowed out Cook Point 18: the absence metric had been spending down the operators’ familiarity with failure for years, and AI removes the last of the slack that hid the bill.
Under AI, Pole B leaders actively study the augmented system’s failure modes while the dashboard is green. They use red-team exercises, deliberate failure injection, and post-incident reviews that ask what did the AI miss, and what would we miss if the AI weren’t here? Bainbridge adds another demand: preserve human access to the patterns automation has displaced, so human monitoring remains real. Pole A leaders won’t run those exercises because the dashboard is green. Pole B leaders run them for the same reason.
A representative case
Suppose a large engineering organisation ends the year with no reportable incidents at all. The leadership team celebrates. A senior engineer privately raises a concern: the team has been working around several known issues for the entire year, and the lack of incidents has more to do with luck and operator skill than with system improvement. The leadership team doesn’t want to hear it. The next year, two of the worked-around issues converge into a major outage that takes down the platform for nine hours.
The post-incident review goes two directions at once. The Pole A direction looks for the proximate cause of the outage. The Pole B direction eventually asks what did the operators know in advance, and why didn’t the organisation hear it?
Hollnagel’s Monitor potential had been missing from the organisation’s investment all along, whatever the dashboards suggested. The operators were monitoring; the organisation wasn’t. The Pole B answer reorganises how near misses get surfaced and resourced. The Pole A answer adds more monitoring of the wrong kind: more dashboards, more alerts. Both moves get made. Only one addresses the cultural condition that produced the year of unsurfaced workarounds.
For engineering leaders, the year of zero incidents is the year to invest in Hollnagel’s Anticipate potential, provided you first run the test that says which kind of quiet year you had. Ask the operators what they worked around. A system whose operators report nothing may be genuinely well designed; a system whose operators have been working around known issues all year has been spending down its slack, and the incident count cannot tell the two apart. Snook, in the same chapter as the F-111 near miss: “Just because it isn’t broken, doesn’t necessarily mean that it isn’t breaking.” Use the quiet period to find out what the operators’ invisible adjustments are holding together. Pole A treats the quiet as success; Pole B treats it as the window during which Pole B investment is cheapest.
The diagnostic move
Three questions for last quarter’s safety review:
- Which pole was I claiming? Did I treat safety as the absence of bad outcomes or as the presence of adaptive capacity?
- Which pole would the actual investment show? Where did the engineering budget you steward actually go: into incident prevention or into resilience capacity? And across Hollnagel’s four potentials, which got the investment? Respond and Learn almost always do. Monitor and Anticipate almost always don’t. Is that distribution the one this system needs?
- Which pole does the system actually require? In any complex sociotechnical system, Pole B is the under-funded half.
The three answers matter most when they diverge.
The exercise
Run a near-miss audit on the last quarter. Collect every near miss, every save, every moment the operators adjusted to keep the system running. Sort them by which of Hollnagel’s four potentials actually did the catching: Respond (the team handled it well when it arrived), Monitor (the team saw it developing), Learn (the team had been here before and knew what to do), Anticipate (the team had imagined this further out and was ready). Most saves involve more than one, so credit the earliest in the chain: if the team saw it developing, score it Monitor, whatever Respond did afterwards. Sorted that way, the distribution shows how early you catch things. The fat buckets are the potentials the organisation has built, whether or not it meant to. The thin ones are where the next near miss will land. A fat Respond bucket is its own warning: the system is catching things at the last possible moment. A second variant for teams that have a strong retrospective practice already: run a Safety-II audit. What’s going right that the team doesn’t have visibility into? Where do operators tell you things work despite the procedure, not because of it? Those answers show where the Monitor potential needs investment, and they are the answers Pole A dashboards systematically miss.
Going upstream
In-text: Erik Hollnagel, Safety-I and Safety-II, and the operational companion Safety-II in Practice (Chapters 4 and 5 for the four resilience potentials and the Resilience Assessment Grid). Richard Cook, How Complex Systems Fail (points 16, 17, 18). Scott Snook, Friendly Fire, on the 1992 F-111 near miss (Chapter 7, the conclusion).
Also touched: The reporting-culture and just-culture conditions that surface near misses are James Reason’s informed culture subcultures, worked at length in the deviation instalment (chapter 14).
Go deeper: The sources under chapter 14—Sidney Dekker, Safety Differently and John Allspaw, How Your Systems Keep Running Day After Day—both teach this dichotomy at the metrics-and-investment level too. The above-the-line / below-the-line frame is from Cook, Allspaw, Woods et al., STELLA Report (2017), unpacked in chapter 14. For pre-accident-investigation methodology: Todd Conklin, Pre-Accident Investigations.
I work with engineering leaders on exactly this kind of paradigm work, the deeper the better. If it’s live for you, I’m happy to talk: schedule a 30-minute virtual coffee at hi.chrisgagne.com.
Some book links here are Amazon affiliate links; if you buy through them I may earn a small commission, at no cost to you.

Leave A Comment