Category: Operating Models & Flow

How software organisations actually deliver: team-of-teams structure, flow economics, cost of delay, and the constraints that decide throughput. Reading the system, not blaming the team.

  • Compromise vs integration

    Compromise vs integration

    You split the difference and called it fair. Next quarter, the same argument is back. What brought it back?


    Chapter 19 asked who should run the decision process. Even a skilled facilitator can reach an agreement that some participants were afraid to question. Before looking for a solution, find out who can safely say what they want.


    The two poles

    Pole A: settle within the available options. Allocate the scarce resource or negotiate a compromise, acknowledging what each party will give up.

    Pole B: change the available options. Test the assumptions that make the two positions collide and look for a solution that meets both underlying needs. The parties must be able to question whether it does.

    Where Pole A is right

    When the resource is genuinely fixed, dividing it is the work: this cycle’s bonus pool, one open VP seat, the headcount available before a deadline. You can still question whether those constraints hold, but a decision may be needed before they can change. Mary Parker Follett, a management thinker writing in the 1920s, argued for integration: finding a solution that meets both parties’ needs. She also recognised that this wouldn’t be possible in every conflict.

    Sometimes the search for an integration costs more than it is likely to save, and a quick fair split, openly made, is worth the time it saves. A clean division can also protect a working relationship better than an optimal solution reached through a bruising process.

    Where Pole B is right

    Pole B is right when the underlying needs can both be met, even if the stated positions appear incompatible. In Follett’s example from the Harvard Library, she wanted to avoid a draught and another reader wanted more air. They opened a window in the empty next room. Both got what they wanted without either having to give something up.

    A compromise that keeps re-opening is worth examining. An unmet need may be bringing it back. The circumstances may also have changed, or the settlement may have been temporary from the start. Recurrence is a reason to investigate.

    In decisions

    Pole A leaders make and explain an allocation or negotiate a split. Pole B leaders look for a change that meets both needs.

    Take a recurring settlement to whoever has authority over the disputed decision. That may be the compensation owner, the budget owner, or the CEO. Ask what each side actually needed and what both were assuming had to be true for these to be the only options. Then question the allocation of authority itself. Follett argues that authority belongs to the job; does the person making this call have the knowledge and responsibility the work requires?

    The sentence your CEO can carry to the board: “We need to understand why these settlements keep coming back before we negotiate them again.”

    Follett’s three ways

    In “Constructive Conflict,” a lecture she gave in 1925, Follett asks the reader to treat conflict as neither good nor bad, “not as warfare, but as the appearance of difference, difference of opinions, of interests. For that is what conflict means⁠—difference.” Difference is unavoidable, so the task is to use it, the way a mechanical engineer both fights friction and runs the belts on it.

    “There are three main ways of dealing with conflict: domination, compromise and integration.”

    Domination is one side winning. Easy in the moment, she says, and not usually successful in the long run. Compromise is the accepted way of making peace by giving something up. Follett argues that the curtailed desire can bring the conflict back. Integration meets both needs: “when two desires are integrated, that means that a solution has been found in which both desires have found a place, that neither side has had to sacrifice anything.” “Compromise does not create,” she writes, “it deals with what already exists; integration creates something new.”

    Follett begins with disclosure: “The first rule, then, for obtaining integration is to put your cards on the table, face the real issue, uncover the conflict, bring the whole thing into the open.” Then you break the demand into its parts, distinguish the declared motive from the real one, and look for a solution that serves both. Before asking people to put their cards on the table, I want to know what it could cost them.

    Artificial harmony

    A team can also avoid addressing the conflict at all. Patrick Lencioni calls the apparent agreement artificial harmony. It makes the work harder because people have to discover what their colleagues think after the meeting has ended.

    Lencioni describes teams resorting to “veiled discussions and guarded comments” when they can’t argue openly. They feign agreement in the meeting and carry the disagreement into hallway conversations and personal attacks. A meeting with no visible disagreement and a car park full of it afterwards is a reason to ask what people couldn’t say in the room.

    The power question

    Who at the table can safely state a real need, refuse, appeal, or walk away? These are different capacities, and each needs a separate check.

    Stating a need takes safety. Amy Edmondson’s work on psychological safety examines whether people can take interpersonal risks, including admitting a weakness or naming a problem. Lencioni’s vulnerability-based trust matters here too. Neither, by itself, gives an employee authority to refuse a decision.

    Refusing takes standing. Chris Voss, the FBI hostage negotiator, writes that “‘No’ is the start of the negotiation, not the end of it.” His “No” expresses the autonomy of a party inside a negotiation. An employee also needs to know what happens if they refuse their manager. A negotiation technique can’t supply that protection.

    Appealing takes a route that the other party doesn’t control. In her paper on responsibility, Follett challenges the illusion of final authority and insists that “authority belongs to the job and stays with the job.” An appeal that runs back only to the person you are appealing to isn’t an appeal.

    Walking away takes material independence. Who can afford to leave this negotiation, and what would leaving cost them?

    Ask who can fire each participant, cut their bonus, or deny their promotion. Who controls the minutes, the information, and the appeal? Barry Oshry built the Power Lab, a residential simulation with built-in differences in power and resources, to study how position shapes what people see. When two parties (his Ends) look to a third (the Middle) to move their competing agendas, “Ends become decreasingly responsible for resolving their own issues and conflicts, while Middle becomes increasingly responsible for resolving these.” Look especially carefully at whoever sits in that middle, and at whose needs are missing from the discussion.

    Follett describes power-with as “a jointly developed power, a co-active, not a coercive power.” She argues that “genuine power is capacity” and that “you cannot confer power, because power is the blossoming of experience.” I take this to mean that integration can also strengthen people’s ability to participate in later decisions. Unequal authority doesn’t make every agreement coerced, but it does require us to examine what an apparent yes means. Someone may be unable to leave a job and still reach an agreement that meets their needs. If they can’t question the proposal without risking their livelihood, we need to address that risk before relying on their assent. An allocation requires the same care.

    When a party attacks the conditions

    Threats or retaliation can make a negotiation unsafe to continue. In that situation, protect people’s ability to participate before asking them to seek a joint solution.

    Karl Popper’s paradox of tolerance, in The Open Society and Its Enemies, helps me think about this limit. He warned that unlimited tolerance could destroy a tolerant society, while cautioning against suppressing views that could still be challenged through rational argument and public opinion. His discussion concerns the defence of a society. Applying it to a leadership meeting is my analogy.

    A colleague remaining unconvinced is insufficient evidence of bad faith. I’d ask what evidence could change their position and make my own answer available too. If they refuse to name any conditions for reconsideration, we can record that refusal and ask why. We still have to distinguish an inference about their motives from what they actually did.

    Where the conduct prevents others from participating, name it and hear the response. If someone must be excluded from part of a decision, specify the restriction, its duration, and how it can be challenged. Review needs to sit with someone independent of the disputed conduct; a manager further up the same chain may not provide that. If no such route exists, acknowledge the limitation and seek one.

    “They are acting in bad faith” is an easy label for a leader to abuse. Excluding someone is an exercise of power with consequences for them. I’d want the judgement open to challenge and the person making it to remain willing to reverse it.

    Goldratt’s cloud

    The best instrument I know for testing the conflict is Eliyahu Goldratt’s Evaporating Cloud, from his 1994 business novel It’s Not Luck. It is a diagram of five parts: a shared objective, two requirements the parties believe necessary to achieve it, and two prerequisites that appear to conflict. The assumptions connecting those parts are what you test.

    When a colleague admires the cloud as a presentation technique, the protagonist answers: “this technique claims that you should not attempt to strive for a compromise. It advocates examining the assumptions under the arrows in order to break the conflict.” The practical advice is to “concentrate on the arrow that irritates you the most.” Read it aloud as a full sentence: “in order to have X we must have Y, because,” and finish the “because.” You now have an assumption you can challenge.

    Goldratt teaches the tool through a teenager’s curfew. His character wants his daughter home before ten for her safety; she wants to stay out later to be accepted by her friends. The assumption connecting the early curfew to safety is that coming home late is itself the danger. Once they look at how she’ll get home, she asks him for a ride, and he agrees. Both needs can be met. The cloud helps them test what each has assumed about the conflict, including assumptions they brought to it themselves.

    When the opposition is real

    Sometimes no change you can make will give both parties what they need. Calling that outcome a win-win conceals the loss. A fair allocation may be the best available decision.

    The person making the allocation needs legitimate authority over it. They should disclose the constraints and criteria, hear objections before the call, and explain who will go without what they wanted. Those affected need a route to challenge the decision. Candidly owning an allocation doesn’t excuse an abuse of authority; the person deciding must also be answerable for how they used it.

    You can also settle provisionally while you keep looking for an integration. Check whether the formulation is poor, whether someone affected is missing, or whether the deadline arrived before a solution did. Follett stresses the intelligence and inventiveness integration requires. Failing to find one isn’t proof that none exists. Record what remains unresolved and agree when to revisit it.

    Voss’s “no deal is better than a bad deal” challenges the reflex to split the difference. His approach may produce a more favourable bargain or a decision to walk away. Goldratt’s cloud asks whether both underlying needs can be met by changing an assumption. I’d choose between these approaches according to the conflict and the parties’ ability to take part.

    The AI overlay

    In an AI deployment, the interests to reconcile belong to the people and institutions involved. A model’s output may supply an argument worth examining, but the model isn’t another stakeholder whose needs the settlement must satisfy.

    Whose needs does the tool serve, and whose work does it now read or direct? Name who chooses the vendor, writes the policy, and decides when an output becomes an action. Also ask who bears the consequences of a mistake. “The model decided” is domination hiding behind a machine. Apply all four questions to the people affected: can they state a need, refuse a proposed use, appeal an outcome, and walk away? What would each choice cost them?

    A representative case

    Consider two engineering leaders disputing a shared platform team’s roadmap. One needs a reliability standard held. The other needs a launch date met. The platform team has a fixed number of weeks, so the decision appears to require a trade-off. In Oshry’s terms, the platform team is the Middle, and both leaders are about to hand it their conflict.

    Suppose you declare reliability “a standing standard” and announce that both needs are met, while assigning the reliability work to people outside the negotiation. The work didn’t stop eating weeks; it moved onto on-call rotations, maintenance, and the product teams who now inherit the acceptance criteria and never got a vote. Those people need to be part of the decision before you can claim to have met everyone’s needs.

    Smaller batches and progressive delivery are candidates to test. Forsgren, Humble and Kim found in Accelerate that high-performing teams achieved both speed and stability. That gives these leaders reason to question the assumed trade-off, but it doesn’t establish that a change in delivery method will meet this launch date and this reliability standard.

    Suppose the cloud exposes a belief that meeting the launch date requires releasing every planned feature together. Test that with the launch owner: what must be available on that date, and which capabilities can follow? Then ask the engineers whether the smaller release can meet the reliability standard within the available time, including the investment in changing delivery practices. Integration depends on both answers, with the people doing the work able to contest the estimates. If the launch genuinely requires every feature, or the reliability work still exceeds the time available, the proposed solution has failed the test.

    Next quarter, suppose the same two leaders need the same senior engineer for work that can’t be shared or resequenced before their deadlines. Check those constraints with them before making the allocation. If they hold, the decision-maker must hear both leaders and explain the criteria for the call. The leader who loses needs to know what work will go undone and how to challenge the decision. A date to revisit it next cycle doesn’t replace a chance to object now.

    The diagnostic move

    Three questions for last week’s contested settlement.

    Which pole was I claiming? Did I describe the outcome as something new that both sides won, or as a fair split, or as a firm call I made?

    Which pole did the outcome actually show? Did both parties get what they came for, or did someone give something up? Who bore the loss, and could they have safely objected? An agreement staying unchallenged tells you little if challenging it would cost someone their job.

    Which conflict is this really? Test the premise that makes the positions incompatible. If the constraint holds, make and explain the allocation. If you have run out of time to investigate, record a provisional settlement. Where conduct prevents people from participating, protect the process before continuing. Recurrence tells you to look again. It doesn’t tell you that you misjudged.

    The exercise this week

    Take one live conflict on your team that keeps re-opening. For each party, check the ability to state a need, refuse, appeal, and leave. Ask what they risk by disagreeing. Establish a protection for any risk that could prevent a straight answer. If the manager controls both the decision and the minutes, for example, arrange for participants to record objections in their own words and have them reviewed independently.

    Then write the cloud. State the shared objective, each side’s requirement, and the two prerequisites in conflict. Complete the “because” sentence under each arrow and test one assumption. Ask whether the proposed change meets both needs, who must do the work, and whether those people can object. If the test fails, settle provisionally and keep looking, or explain why the remaining constraint requires an allocation.

    You may leave the meeting with a tested integration, an acknowledged allocation, or a provisional settlement. Record which it is and what happens next.

    Going upstream

    Start with Mary Parker Follett’s Dynamic Administration, particularly “Constructive Conflict” and her papers on power and responsibility. Eliyahu Goldratt’s It’s Not Luck gives you the Evaporating Cloud and worked examples of testing the assumptions beneath a conflict.

    Patrick Lencioni’s The Advantage and The Five Dysfunctions of a Team discuss artificial harmony and fear of conflict. Amy Edmondson’s The Fearless Organization examines psychological safety. Barry Oshry’s Seeing Systems shows how organisational positions shape people’s experience and behaviour. Chris Voss’s Never Split the Difference makes the case against reflexive compromise. Nicole Forsgren, Jez Humble and Gene Kim’s Accelerate reports the research on software delivery performance.

    For further reading, Jeffrey Pfeffer’s Power examines organisational power. Karl Popper’s The Open Society and Its Enemies contains the paradox of tolerance; his discussion of defending a tolerant society needs care when applied to workplace decisions.


    I work with engineering leaders on exactly this kind of paradigm work, the deeper the better. If it’s live for you, I’m happy to talk: schedule a 30-minute virtual coffee at hi.chrisgagne.com.

    Some book links here are Amazon affiliate links; if you buy through them I may earn a small commission, at no cost to you.


  • Prioritising a portfolio

    Prioritising a portfolio

    Weighted Shortest Job First is the right instinct: sequence work so the things whose delay costs the most, relative to their size, go first. I learned WSJF inside SAFe, which scores its components in relative, modified-Fibonacci points because dollar figures are rare; the economics underneath are Don Reinertsen’s (The Principles of Product Development Flow). When a client has believable, dollarised cost-of-delay numbers, I’ll sequence on them all day. Most portfolios I meet don’t have those numbers. Cost of delay exists for the two or three epics somebody fought about, is contested for a dozen more, and is guesswork past that. You can stall for a quarter trying to dollarise the whole list, or you can get a defensible forced-rank order out of the preferences the organisation already holds.

    I’ve run the SAFe-style scoring version myself, coaching a prior client’s leadership through scoring value, time criticality, risk, and job size in Fibonacci points in a shared spreadsheet: eight hours of discussion over four days. One product leader had described the state before it as needing to address five large initiatives at once. The exercise landed (I suspect it gave the organisation its first formal Product Backlog, and one held with broad support), but eight hours of invented numbers is a heavy lift to repeat. Forced choice keeps the agreement and drops the arithmetic.

    My primary mechanism for the second path is adaptive forced choice: paired comparisons, the preference-elicitation move conjoint studies in market research are built on. Put two epics on the screen, ask which one the organisation would rather have, and let many small, easy judgements assemble a ranking that no single big judgement could produce. Where a SAFe-style scoring session asks the room to agree that time criticality is an 8 rather than a 13, forced choice asks a smaller question⁠—this epic or that one⁠—and nobody has to invent a number. I run it through two channels at once. Stakeholders vote in the room, in Delphi rounds, with Plickers cards. Engineers work directly in Prioneer.io, the same engine that serves the room its comparisons.

    The room

    Prioneer drives the screen, serving one comparison at a time⁠—”Epic F vs Epic R”⁠—with four answers: A, strongly prefer left; B, moderately prefer left; C, moderately prefer right; D, strongly prefer right. Everyone holds up a Plickers card, a paper square whose rotation encodes the answer. Three properties make the cards better than raised hands or a polling app for this job. A neighbour can’t read a Plickers card at a glance, so the first vote lands before anyone can anchor on the boss’s arm. Nobody needs a phone or a laptop to vote⁠—the only device in the room is mine⁠—so people stay present instead of drifting into their screens. And I scan the whole room in a few seconds, so the tally is on the screen while the choice is still warm.

    Then the Delphi move: where the room splits, we hear the strongest voice on each side and vote the same pair again. (Delphi is the RAND-developed technique of iterated judgement with controlled feedback; mine is a lighter, in-room version.) In my experience a second vote after two minutes of argument moves more minds than an hour of open discussion, because the argument has a decision to land on. Group preferences can cycle⁠—a room may prefer F over P, P over R, and R over F⁠—and I leave those to Prioneer, whose ranking model reconciles the full answer set rather than trusting any single chain of inference.

    The engineers

    Engineers skip the cards and make their trade-offs straight in Prioneer, whose pairwise ranking is adaptive: once it knows F beats P and P beats R, transitivity means it stops asking about F and R, so nobody grinds through every possible pair. They rank the same epics on complexity, uncertainty, and effort: the denominator side of the WSJF ratio, judged by the people who’ll carry the work.

    What comes out

    A full forced-rank order of the portfolio, and, more usefully, the disagreements. Where the stakeholder ranking and the engineering ranking part company, you’re looking at WSJF’s numerator and denominator arguing with each other, and that argument is the agenda for the next working session.

    A perfect ordering was never the goal; what you need is enough stakeholder alignment to limit work in progress. Swapping two adjacent epics in rank costs little. Running twelve at once costs a context-switching tax on all twelve. Larman and Vodde call that the Context-Switching Vampire, fed by well-intended, perfectly reasonable requests from every direction, and their structural cure in LeSS is a single source of work for the teams. The forced-choice pass buys the alignment that makes a single source survivable: once the room has watched the order assemble out of its own votes, the organisation can leave the tail of the list unstarted without relitigating it every week.

    At one client, a public B2B SaaS company, I ran this with 30+ stakeholders including C-level leadership, and it reset the priorities of the entire epic portfolio. The final sequence was based heavily on the quantitative results, with the stakeholder-versus-engineering gaps taken to working sessions. The decisions were theirs; the mechanism made the preferences visible enough to decide on. Work in process fell, throughput rose, and we realised value sooner.

    When you do have the dollars

    Use them. At another client I built deterministic cost-of-delay sequencing into the delivery tooling itself⁠—Delivery Intelligence—with forecast date ranges at a stated confidence (P50–P95); where the business states what an epic is worth, the tool can put an estimated dollar figure on the delay. The forced-choice pass gets a portfolio moving while the data matures, and it builds the shared preference structure that makes the later dollar arguments tractable.

  • Expert vs facilitator

    Expert vs facilitator

    You’re the last word on the technical call. You’re also the bottleneck on every material technology decision.


    Chapter 18 looked at the assumptions a leader holds about workers. The role the leader takes comes next, because the expert stance reverses those Theory Y intentions at the moment of the call.


    The two poles

    Pole A: expert. The leader is the most senior technical authority. Problems escalate up to where the expertise is.

    Pole B: facilitator. The leader’s role is to make others’ decisions possible. Expertise lives at the work.

    Where Pole A is right

    When the leader is genuinely the deepest expert and the call is technical. The founding CTO making the last call on a foundational architecture decision. The chief surgeon overruling on the table. The lead investigator on a specific case where the situation depends on their direct expertise. Pole A is right in those contexts.

    Where Pole B is right

    Almost every senior leadership context outside narrow technical-founder situations. By the time leaders reach senior roles, their expertise is usually a layer or two removed from where the work is being done. When it is that far removed, deciding anyway usually produces worse calls than the team would have made unaided.

    In decisions

    Pole A leaders decide contested technical calls. Pole B leaders convene the people who know and back the call they make.

    Take the last contested technical decision that routed to you and ask where the fix sits: how should decisions like this route, and what would have to be true for them to stop routing to me? Most of it is inside your own function, and there it isn’t an ask at all: I intend to push architecture calls to the team closest to the system, because that is where the expertise lives; unless you see something I don’t, it starts with the next decision.

    What belongs to the CEO is the part only the CEO can move: holding the function heads accountable for building the capability rather than escalating it, and authorising the change when the routing crosses your boundary. Take that part to them in the same open form: what would have to be true for the function heads to build this capability instead of escalating it? The authorisation is often a small word change, what Marquet calls a genetic-code move: changing the words the organisation runs on.

    The sentence your CEO can carry to the board: “Every decision that escalates to us loses some of the context it needed on the way up.”

    Schön’s reflective practitioner

    I take Donald Schön’s The Reflective Practitioner as the canonical critique of technical rationality: the assumption that expertise consists of applying codified knowledge to standardisable problems. Schön found, across many professional practices, that real expert work is situational. The cases that count are the ones that don’t fit the textbook. The expert’s actual capability is reflection-in-action: thinking inside the practice and adjusting in real time to build a unique response to a unique situation.

    Schön’s critique cuts hardest at the expert-as-authority model. If real expertise is situational, the senior person who decides from a distance is working with less of the situation’s live context than the practitioner who is in it. The senior person’s broader experience matters for some questions and is irrelevant for others. Pole B leaders learn to tell those questions apart.

    If real expertise is situational, the senior person who decides from a distance is working with less of the situation’s live context than the practitioner who is in it.

    Schön goes further with what he calls the reflective contract: a relationship in which the practitioner can admit uncertainty. The client gets a partner in inquiry rather than a vending machine for answers. The leader promoted into senior expertise can step into that contract with their team. Most of the leaders I’ve worked with don’t. The expert posture is what got them promoted, and the switch is uncomfortable in exactly the ways promotions aren’t.

    Block on the consultant-leader collusion

    Peter Block names the pattern that keeps Pole A in place. The content level is the technical question, the framework, the deliverable. The affective level is the relational dynamics, the power and trust in the room, and the question of who the leader takes themselves to be. For years I worked the content level too, as chapter 1 says; what I would add here is that the preference is mutual, and that is what makes it stick. Block’s word for what it produces is collusion, and calling the content-over-affect preference itself collusion is my generalisation of his pattern. The consultant gets to look smart; the leader gets to look in charge; the conditions that produced the problem stay untouched.

    The leader has to make the same Pole B shift in the consultant relationship. Here, the leader sits in the consultant’s chair vis-à-vis the team. Pole A says the leader is the deepest expert present. Pole B says the leader’s role is to make the team’s existing expertise legible to itself.

    In your last contested team decision, who was at the centre: you, with the answer, or the team, with the problem they own? Putting yourself there closes the immediate decision faster and takes the call out of the team’s hands, which makes the next contested call more likely to route to you.

    The Scheins’ four kinds of inquiry

    The Scheins’ Humble Inquiry grounded the expert-vs-facilitator argument for me for years. The same recordings chapter 5 drew on show a second pattern in the early sessions: I would explain the framework, parts theory, polyvagal, whatever fitted, while the client was still in the middle of a felt sense. What was missing was the confrontive question that would have kept them in it. Humble inquiry is the canonical question Pole B asks: the genuine question to which the asker doesn’t know the answer. It’s the right starting place, and the rest of their framework does other work.

    The Scheins name four types of inquiry. Humble inquiry is the question the asker has no answer to. Diagnostic inquiry steers the conversation toward what the asker thinks is relevant: causes, feelings, actions taken. Confrontive inquiry goes further and inserts the asker’s own idea in the form of a question. Process-oriented inquiry is the question about how the conversation itself is going.

    The four aren’t a hierarchy. The Pole B leader carries all four and picks among them as the situation calls. Their own counsel is to blend the helping forms with humble inquiry according to the needs of the situation and to stay conscious of the switch; their warning is that the helping forms take charge of the conversation and slide toward telling.

    Ask: which kind of inquiry am I in right now, and which would best serve the situation?

    Weinberg’s secrets

    Jerry Weinberg names the working discipline. The 10% rule from chapter 1 applies here too: never promise more than 10% improvement. My read is that the rule holds because anything more requires moving from content to affective level, which the leader will resist. Weinberg’s real prize is the credibility of the naive question. His Golden Lock names the trap⁠—I’d like to learn something new, but what I already know pays too well—and the seniority version is my extension: the senior person who has stopped being able to ask basic questions because appearing not to know would threaten their standing. The Pole B leader recovers the capacity to ask the question the team needs asked, even when the asking exposes that they don’t know the answer. NUMMI’s 1984 team member handbook, from the plant in chapter 18, told the people on the line the same thing: “If you accept an explanation without question you may have lost the chance to understand. You must learn to say ‘I don’t understand.’” It is a harder sentence for the person who is supposed to know.

    The question costs status. Ask it anyway.

    O’Neill⁠—backbone and heart

    Mary Beth O’Neill draws the distinction sharper than Block does. She names two operating models for the helper: the rescue model, which puts the helper at the centre (the helper has the answer, the client has the problem, the helper fixes it), and the client-responsibility model, which puts the client at the centre (the client owns the problem, the helper supports their own work on it). Block’s collusion is the rescue model, and expertise is what excuses it. The expert-leader running the rescue model uses role-authority to take responsibility for the call. The team learns not to bring discretion; the rescue reproduces its own need.

    Her title names the stance the Pole B shift asks for: backbone and heart, knowing your position and stating it clearly while staying engaged in the relationship, so the team feels met rather than processed. The harder part is holding it even when the team is waiting for the leader to decide.

    Joiner’s Catalyst stage⁠—and the levels around it

    Bill Joiner and Stephen Josephs name the developmental stage the Pole B shift belongs to. Chapter 12 met their ladder from the other end; the transition that matters here is Achiever to Catalyst.

    The transition from Achiever to Catalyst is the change this chapter describes. The Achiever makes decisions well; the Catalyst develops the people who will make them. That change is uncomfortable because Achiever skills are rewarded and visible, while Catalyst skills stay invisible until the team starts producing decisions the leader couldn’t have made themselves.

    The Achiever makes decisions well; the Catalyst develops the people who will make them.

    The ladder matters whole, not collapsed to Catalyst. The Expert stage is apt when the leader is the deepest expert on the specific question: the founding CTO on a deep technical call, the chief surgeon on the table. The error is generalising that posture to situations that need Catalyst or Co-Creator. The Synergist can move between stages as the situation fits without turning any stage into an identity.

    A representative case

    This is a composite of leaders I have sat across from. A senior architect gets promoted to VP of Engineering, deserved on contribution, and it is a problem in waiting. Six months in, the VP is making most of the architecture decisions for the engineering organisation, and most of their engineering managers have stopped bringing recommendations. They were promoted because they were Experts or Achievers, and nobody explained the need for them to become Catalysts.

    The Pole B shift is slow. The first move is Marquet’s I intend to: the VP makes every engineering manager state intent rather than ask permission. Then the VP has to stop having an opinion on every technical call and ask the team for its recommendation without offering their own. Hardest of all, the VP has to publicly back the team’s call when it differs from what they would have done, not in the private moments where it is politically safe but in the all-hands meetings where their own credibility is on the line.

    What changes if it lands is not the quality of any single call. Some will be better than the VP would have made and some worse. What changes is that the calls get made without the VP in the room, and the company can operate without them, which it could not before. The VP also gets their evenings back. Marquet names the cost of the other way in his own telling of the I intend to shift: if you are always the answer man, you can never go home and eat dinner.

    The VP doesn’t stop being technically deep; they stop putting the technical depth into the seat the team is supposed to occupy.

    From facilitator to practitioner

    This chapter sets up the practitioner figure. The expert is the helper who knows the answer: Pole A, the technical-rationality position. The facilitator is the helper who makes the team’s answer possible: Pole B, the Catalyst-stage position. The practitioner figure that chapter 22 will name is the deepest version of the Pole B shift: the helper who can stay steady with a leader while their old sense of the job comes apart, without pushing it to a tidy resolution, when the situation calls for more than facilitation.

    All three are asked to do the same thing. Step out of the content the team is responsible for. Bring the developed self O’Neill calls signature presence, the helper’s own self as the instrument that makes the team’s work possible. Hold the conditions and leave the decisions with the team.

    The diagnostic move

    Three questions for last Tuesday’s contested call.

    • Which pole was I claiming? Did I describe my role as expert or as facilitator?
    • Which pole would my actual behaviour show? Did I make the call myself or convene the people who knew?
    • Which pole does the situation actually require? Unless I am genuinely the deepest expert on the specific question, Pole B.

    The exercise this week

    Run the convening exercise once this week. Pick a decision you would normally make as the technical authority. Set out the problem to the team in 20 minutes. Ask the team to recommend. Listen, without correcting, for the full conversation. Back the recommendation publicly, even if it isn’t what you would have done.

    In my experience, the team reads your willingness to back its recommendation more than the recommendation itself. The exercise is one decision. For the next 30 days, track how many contested technical calls the team resolves without you and how many still route back.

    If the recommendation is genuinely wrong on the merits⁠—rare in my experience, but it happens⁠—start by asking the team what evidence would change their recommendation, and share the evidence you have that they don’t. If the evidence persuades them, the team makes the better call. If it doesn’t, you have a real disagreement. In a Pole B decision, the conversation continues and the call remains with the team. In the narrow Pole A case, the leader makes the call and owns why.

    Going upstream

    Watch. Edgar Schein, Humble Inquiry lecture, a short talk for leaders on his own channel.

    In-text: the two main works named in the body. Donald Schön, The Reflective Practitioner, on technical rationality and the reflective contract. Edgar Schein and Peter Schein, Humble Inquiry (2nd ed.), on the four kinds of inquiry, not just the humble one.

    Also touched: the helper models named in the body, with full references here. Peter Block, Flawless Consulting, on contracting at the affective level and the forms of consultant-client collusion: pretending the organisation is rational rather than political, skipping the touchy contracting subjects, rushing to solutions before the problem is owned. Jerry Weinberg, Secrets of Consulting, for the named laws, the 10% rule, and the Golden Lock. Mary Beth O’Neill, Executive Coaching with Backbone and Heart, on signature presence (the coach’s own developed self as the instrument), the rescue versus client-responsibility model (a contrast she credits to her colleague Rob Schachter), and the four-phase coaching methodology of Contracting, Planning, Live-Action, and Debriefing. Bill Joiner and Stephen Josephs, Leadership Agility, on the five-stage developmental model⁠—Expert, Achiever, Catalyst, Co-Creator, Synergist⁠—worked at full strength, with the percentages, in the socialised-versus-self-authoring instalment. The NUMMI line is from the plant’s own 1984 team member handbook, page 53, scanned from the Walter P. Reuther Library archive. Marquet’s answer-man line is from his 2013 Inno-Versity talk, Greatness, and not from either of his books.

    Go deeper: Jerry Weinberg, More Secrets of Consulting: The Consultant’s Tool Kit, for the catalogue of named devices in his Consultant’s Self-Esteem Tool Kit, six of them inherited from his teacher Virginia Satir, the rest his own and his colleagues’. My own earlier pass at this argument is the 2018 Scrum training Agile Fundamentals: Going from Expert Individual Contributor to Humble Leader, which runs the same shift through Cynefin and Toyota’s NUMMI handbook.


    I work with engineering leaders on exactly this kind of paradigm work, the deeper the better. If it’s live for you, I’m happy to talk: schedule a 30-minute virtual coffee at hi.chrisgagne.com.

    Some book links here are Amazon affiliate links; if you buy through them I may earn a small commission, at no cost to you.


  • Theory X vs Theory Y

    Theory X vs Theory Y

    Your stated values are Theory Y. The approval thresholds, access controls, and review gates inside your function are Theory X. Your team reads the org design, not the values.


    Leader-leader from chapter 17 can’t run on Theory X assumptions about workers. Unless those assumptions change, the structure pulls the protocol back to leader-follower at the next reorg.


    The two poles

    Pole A: Theory X. Design control structures around the worst-case employee. Oversight as the default. Trust is a reward earned through visible compliance.

    Pole B: Theory Y. Design control structures around the average employee. Trust the system to handle the exceptions. Remove the obstacles, and let people do good work.

    Douglas McGregor drew the distinction in The Human Side of Enterprise in 1960. More than 60 years on, most organisations still operate on Theory X assumptions even when leadership quotes Theory Y.

    The web of our own weaving

    McGregor’s deepest claim wasn’t that Theory X is wrong about people. It was that managerial strategy produces the workers it predicts. He called it a web of our own weaving. The leader who treats workers as untrustworthy designs controls that punish discretion. The workers learn that discretion is punished, so they stop bringing it. The leader observes the absence of discretion, concludes the original assumption was correct, and tightens the controls. The system has made its own evidence.

    This turns the question from a values question⁠—do you trust your people?—into a feedback-loop question: what is your org design⁠—strategy, processes, and rewards as much as reporting lines⁠—producing in the people who run inside it? As a values question, Theory X / Theory Y is the kind of debate the leader has every five years at an off-site and forgets by the following Monday. As a feedback-loop diagnosis, it is something the leader can read and action immediately.

    Run McGregor’s loop backwards. List the behaviours your control structure is currently producing in your engineers. Then ask whether those behaviours were latent in the people or installed by the structure. In my experience the answer is some of both, and the structural share rises the longer they’ve been inside the organisation.

    Where Pole A is right

    McGregor named two contexts where Theory X structures genuinely fit. Both still hold today.

    First: genuinely high-stakes regulated work where the cost of a single bad actor is catastrophic. This includes nuclear plant control rooms, aviation cockpit checklists, surgical theatre protocols, anti-money-laundering compliance, and pharmaceutical batch release. The structural Theory X here is the right engineering response to predictable work with tightly bounded failure modes.

    Second: routine low-skill, low-discretion work where the task itself doesn’t generate intrinsic motivation and the structure is doing the load. High-turnover transactional roles and certain warehouse or contact-centre tasks fit this description, particularly where the work is the work and capability development sits outside it. Intrinsic motivation isn’t available to carry the load, so the structure has to compensate.

    The error is generalising from these contexts to most knowledge work. Most knowledge work, including cross-functional work, is neither. Roles that depend on intrinsic motivation usually need Pole B, yet many inherit Pole A structures through habit rather than engineering judgement.

    Where Pole B is right

    Theory Y fits wherever the work depends on judgement the worker has and the control structure can’t supply. That is McGregor’s claim about knowledge work, and software-dependent organisations are where it bites hardest; run the Vodde exercise below with your teams and see for yourself.

    A Theory X org design taxes the capacity the work runs on whenever the task is non-routine, the person closest to it knows more than the layer above, and the outcome depends on initiative. The engineer who must seek approval for time to refactor and the team whose every experiment routes through a change-approval board face the same constraint. The structure punishes the initiative the work needs.

    Remove the obstacle and the average person does good work, because the work itself supplies the motivation the controls were trying to manufacture.

    The AI shift shows how much work already depended on human judgement. As AI absorbs the routine and the rule-bound, what stays human is the judgement-heavy, initiative-bearing work where Theory Y is an economic necessity, not an emotional preference. My read is that the judgement-heavy work was always the larger share, and the Theory-X-shaped routine sitting on top of it hid how much. Strip the routine away and most of what remains needs Pole B. The leader still running Theory X org designs across that work is paying a compounding tax on the one capability the AI era requires most: judgement.

    Honour the past⁠—why Theory Y didn’t diffuse cleanly

    Theory Y has had 60-odd years to spread, but it hasn’t. A leader who waves it around as the obvious correct answer still has to reckon with the question McGregor’s successors faced: Why didn’t it diffuse?

    Theory Y adoption has been real but partial, often reversed by leadership succession, and dependent on conditions most companies didn’t sustain. A leader installs a Theory Y organisation design, the design produces results, the leader leaves or gets promoted, and the new leader treats the design as overhead and reverts to Theory X defaults. The workers who had developed inside the Theory Y structure now look like overpaid prima donnas to a Theory X eye.

    A successor can restore the old controls faster than workers learned to use the discretion the first leader gave them. That succession pattern is what Kochan, Orlikowski, and Cutcher-Gershenfeld documented: the pipeline mostly produces Theory X operators, and the Theory Y operator is the one who has to be found.

    The leader who currently runs Theory X org designs isn’t necessarily running them out of bad faith or ignorance of McGregor. They are running them because Theory X is what the system selected for, what their training reinforced, and what their predecessors left behind. The leader’s previous paradigm got them here. The past wasn’t wrong; the past fit its conditions.

    McGregor also named a failure mode for the leader who thinks they’ve adopted Theory Y but hasn’t. Old wine in new bottles: relabelling Theory X with Theory Y vocabulary. He wrote of decentralisation, management by objectives, consultative supervision, and democratic leadership in his own era. The contemporary equivalent is psychological safety as a poster sitting on top of an unchanged escalation pattern, or empowerment as a slogan sitting on top of an unchanged approval workflow. Each is a Theory Y label sitting on top of an unchanged Theory X structure.

    Bas Vodde’s exercise⁠—McGregor’s diagnostic, repackaged

    Bas Vodde, Craig Larman’s collaborator on Large-Scale Scrum (LeSS), teaches an exercise that catches the gap in flight. Draw a horizontal axis. Theory X to the left, Theory Y to the right. Sticky-notes above the line for observed attitudes and behaviours of fellow colleagues. Sticky-notes below the line for the organisation design the team is operating inside. The gap between the two rows is the diagnostic.

    In the Vodde class I sat in on, and in my own work since, the pattern is consistent. Above the line: Theory Y. Below the line: Theory X. The leader’s values are above the line. The structure they run is below the line. Both are sincere. The team reads the structure.

    The exercise works because it externalises what the leader already knows. The gap was visible to the team before the exercise; the exercise makes it visible to the leader. After that, the old account is harder to sustain.

    The exercise isn’t Vodde’s invention. McGregor named the same gap in the original book and called it the lip-service Theory Y problem: the leader who quotes Theory Y while running Theory X is doing what McGregor was already cataloguing in 1960. Vodde’s contribution is the experiential packaging: the sticky-notes, the horizontal axis, the visible gap. McGregor named the diagnostic; Vodde made it teachable in 20 minutes with a pad of sticky notes.

    The operating conditions Theory Y assumes

    Theory Y assumes the work supplies its own motivation: that the worker chooses how to do it, that the work develops skill the worker values, and that it connects to something larger than the worker. Where those conditions hold, intrinsic motivation does most of the work the Theory X structure was trying to do. Where they don’t, no Theory X structure can fully substitute. The structure can extract compliant behaviour; it can’t extract the discretionary effort the work actually needs.

    Theory Y structures don’t cause good performance, they permit it. The good performance was always going to come from the intrinsic motivation. Theory X structures suppress it; Theory Y structures stop suppressing it. From a control frame, Theory Y looks like abdication. From an autonomy frame, it looks like taking the brakes off. The Pole B leader sees it as removing the brakes.

    Theory Y structures don’t cause good performance, they permit it.

    The compliant, externally-motivated worker Theory X assumes and the intrinsically-motivated, autonomy-driven worker Theory Y assumes can be the same person under different structures. The surrounding structure draws out the behaviour. NUMMI is the closest thing to a controlled experiment. When the Fremont plant reopened in 1984 under Toyota’s system, over 85% of the workforce were old hands from GM Fremont, the same people the industry had called the worst workforce in the automobile industry; within three months the cars coming off the line were getting near perfect quality ratings. Bruce Lee, who ran the union local and fought to have them rehired, put the reading plainly: “I believed that it was the system that made it bad, not the people.” This is McGregor’s web of our own weaving, still being woven in approval thresholds and sign-off chains.

    The coupling with leader-leader

    Theory Y is what makes Marquet’s leader-leader possible. Chapter 17’s leader-leader move depends on subordinates having the orientation and motivation to act on intent rather than instruction. After years of Theory X, people have learned not to bring discretion to work. Install leader-leader at that point and they keep escalating decisions the new structure expects them to make. The confusion is learned.

    The two axes are coupled. Pole A on this axis forces Pole A on leader-follower. You can’t run leader-leader on Theory X assumptions; the assumptions defeat the protocol. A Pole B move on either axis is incomplete without the Pole B move on the other.

    The leader-leader chapter’s submarine case turned on exactly this: the structural change could work because the captain trusted the crew to have the orientation Theory Y predicts, and the crew rose to it. The change ran on one set of assumptions about people, held consistently enough that the structure could move.

    Where Theory Y gets misread

    Of the companies I have been inside, most had softened the structures with culture work and kept the underlying assumptions. I have not been inside one that rebuilt the structure around Theory Y assumptions from the ground up.

    Theory Y is not a developmental destination, though it gets read as one. McGregor’s claim was that the assumptions a leader holds about workers produce the workers they predict, and that for most knowledge work those Theory Y assumptions produce a better outcome than Theory X ones.

    McGregor’s account was situational. Pole A fits some situations; Pole B fits more situations than most leaders think. The developmental reading is a later overlay.

    A representative case

    Picture a mid-sized engineering organisation with a documented value of trust your people. The values slide is shown in every monthly all hands. The actual budget approval process requires four signatures for any spend above $500. Travel requires CEO approval. Time-off requests go through three managers. Code-review policy requires two reviewers for every change, regardless of risk profile.

    A new VP of Engineering runs Vodde’s exercise in their first week. The above-the-line stickies describe colleagues worthy of trust. The below-the-line stickies describe a system of suspicion. The gap is obvious to everyone on the team and has been invisible to everyone in leadership for years.

    The Pole B move takes six months. Budget authority pushes down to team leads for spend below $5K. Travel approval delegates to engineering managers. Time-off requests require only the engineer’s own calendar coordination. Code review reduces to one reviewer for low-risk changes.

    In McGregor’s vocabulary, the web is now working in the other direction. The Theory X structure had been producing Theory X behaviour. The structural change permits Theory Y behaviour. The engineers haven’t changed; the structure has stopped suppressing the discretion they already had.

    The AI overlay

    AI tooling extends both poles. Pole A leaders use AI for monitoring and compliance: automated detection of policy violations, AI-driven approval workflows, predictive surveillance of engineer behaviour. The Pole A AI stack codifies Theory X assumptions into infrastructure. Pole B leaders use AI to extend autonomy: agents that handle the routine approvals so the human chain doesn’t bottleneck, AI assistants that make the intent-contract legible across larger teams. The Pole B AI stack codifies Theory Y assumptions into infrastructure.

    Both poles can use the same tools. Their configuration tells the truth about the leader’s assumptions, so a leader genuinely unsure which pole they’re on can read it off the AI stack they have built.

    Ask the model which way to configure and it leans toward the paradigm you bring it. A leader who believes people need watching will get a confident, well-sourced design for monitoring, approval gates, and surveillance, because the management corpus the model learned from is thick with exactly that. A monitoring-framed prompt tends to return monitoring designs; an autonomy-framed prompt returns autonomy designs. The tool deepens the groove.

    The redesign-question argument the series opened with arrives here too. AI tooling makes paradigm-misfit cheap to deploy. The tooling is the same either way: same models, same licences. What differs is what the deployment produces in the people running inside it, and the web McGregor named now gets industrialised at AI scale, in either direction, faster than the organisation can decide which web it actually wants to weave.

    The diagnostic move

    Three questions for last Tuesday’s structural decision.

    • Which pole was I claiming? Did I say I trust the team or that I needed to maintain controls?
    • Which pole would the actual structure show? If I ran Vodde’s sticky-note exercise tomorrow, what would the gap reveal?
    • Which pole does the work actually require? Most knowledge work is Pole B. The Pole A defaults are leftovers from a paradigm that fit other conditions.

    CTOs and VPEs: run Vodde’s sticky-note exercise one layer up, with your CEO as co-author rather than subject⁠—each of you holding the pen on your own rows⁠—before you assume the gap is yours to close. Above the line, the stated values from the culture page, the values slide, and the all-hands language. Below it, three things the CEO directly controls: budget approval thresholds, headcount sign-off, and strategic hire veto. McGregor’s web works in both directions: if the CEO runs Theory X assumptions above the CTO layer through approval workflows, access controls, and budget signatures, the org design will pull back any Theory Y installation below. The gap is what your engineers read.

    Ask for one small, specific structural change that brings one below-the-line item into line with the stated value. The CEO doesn’t have to redesign the organisation. They have to move one threshold. The sentence your CEO can carry to the board: “Our company runs on the org design, not the values slide.”

    The exercise this week

    The reflection is solo. The exercise is for the team.

    Run Bas Vodde’s sticky-note exercise with your direct reports this week. 20 minutes. Draw a horizontal line on the wall. Theory X to the left, Theory Y to the right. Each person writes three sticky-notes above the line (three behaviours or attitudes they observe in the team) and three sticky-notes below the line (three structures or processes the team is operating inside). Place the stickies on the axis. Stand back and look.

    Where is the gap widest? Which structure is most clearly Theory X? Which Theory X structure is doing genuine load (high-stakes risk surface, low-discretion routine work), and which is just inherited habit? What’s the smallest structural change that would let the team’s actual Theory Y behaviour stop being suppressed?

    Above the line: Theory Y. Below the line: Theory X. Structures that match the value are worth strengthening.

    Going upstream

    In-text: Douglas McGregor, The Human Side of Enterprise: the canonical primary. A web of our own weaving, old wine in new bottles, and the lip-service Theory Y problem all live here.

    Also touched: Bas Vodde, Craig Larman’s LeSS collaborator, for the above-the-line / below-the-line sticky-note exercise. The NUMMI figures and Bruce Lee’s line come from This American Life’s “NUMMI” episode (2015).

    Go deeper: Daniel Pink, Drive, on the operating conditions Theory Y assumes⁠—autonomy, mastery, purpose⁠—and the Type X / Type I distinction; the short version is RSA ANIMATE: Drive (11 min). L. David Marquet, Turn the Ship Around!, covers the Theory Y assumptions behind the leader-leader move developed at length in the leader-leader instalment. For the empirical-diffusion question⁠—why Theory Y reverts at leadership succession⁠—see Kochan, Orlikowski, and Cutcher-Gershenfeld, Beyond McGregor’s Theory Y (MIT Sloan, 2003).


    I work with engineering leaders on exactly this kind of paradigm work, the deeper the better. If it’s live for you, I’m happy to talk: schedule a 30-minute virtual coffee at hi.chrisgagne.com.

    Some book links here are Amazon affiliate links; if you buy through them I may earn a small commission, at no cost to you.


  • Leader-follower vs leader-leader

    Leader-follower vs leader-leader

    You spent Tuesday in approval meetings. You spent Wednesday wondering why your team can’t make a decision without you, and why the function still can’t ship faster. The structure is leader-follower, all the way down.


    Chapter 16 named the discipline of small batches and pull. This chapter asks who holds the authority to decide what gets pulled, because most software organisations route every non-trivial decision through the senior person’s calendar.


    The Tuesday calendar is the visible artefact. The Wednesday frustration is the felt consequence. Between them sits a paradigm so habitual that the people inside it usually can’t see it as a paradigm: authority lives with the senior person at each layer, subordinates seek permission, orders are detailed, and the senior person’s bandwidth is the rate limit of the whole layer below.

    Marquet calls it leader-follower.

    This industrial-management inheritance remains the default. I suspect most software-dependent organisations run it across at least one major function without anyone at the executive level recognising it as a paradigm rather than the way things are.

    The other pole exists. David Marquet, who took command of the USS Santa Fe in 1999, a submarine with the worst retention record in the force, gives it the operational name leader-leader. Authority lives where the information lives. Subordinates state intent. The leader replies Very well. Orders carry purpose and constraints, and the team chooses the action.

    Marquet was the captain. His executive officer, the XO, ran the officers; his chief of the boat, the COB, was the senior enlisted man aboard and ran the chiefs. The crew did the work. The captain’s job was to mandate the structural shift, model the language, and refuse to make the decisions the captain shouldn’t be making.

    In 1998, the year before he arrived, the boat had re-enlisted three sailors out of a crew of 135. In Marquet’s first full year it re-enlisted 36, and by 2001, the last year of his command, it had earned the highest grade on its reactor-operations inspection that anyone had seen. The book’s afterword reads like a command roster: officer after officer of that wardroom went on to command submarines of their own. Santa Fe shifted from one leader and 134 followers to 135 leaders.

    This chapter is about the move at the executive level. Marquet’s captain is the CEO. The XO and COB are the senior leadership team: the CTO, the VPE, the heads of function. The crew is the rest of the organisation.

    An organisation-wide, durable shift usually requires CEO participation.

    The CTO who tries alone usually produces leader-leader inside one team while the surrounding structure pulls it back at the next reorg or budget cycle. Santa Fe is the exception the In-decisions section leans on, and it took Marquet doing it from the middle.

    The two poles

    Pole A: leader-follower. Authority and decision rights concentrate with the senior person. Subordinates seek permission before acting. The leader spends Tuesday in approval meetings because the structure requires the approvals. Orders specify what to do. Compliance is the metric of execution. The leader’s working day fills with decisions that arrived at the desk because no one else had the authority to make them. The bandwidth of the leader is the rate limit of the team.

    Pole B: leader-leader. Authority and decision rights live where the information lives, which is almost never with the senior person. Subordinates state what they intend to do. The leader communicates purpose and constraints, not action.

    Marquet’s specific protocol is “I intend to…” The officer or sailor states what they’re about to do. The captain says Very well.

    Marquet frames the shift as three replacements: trading pressure and conformity for clearly stated intent, swapping the urge to prove oneself for improving and learning, and giving up the show of certainty for openness and curiosity. Subordinates aren’t waiting to be told. They are thinking, intending, and acting. The leader contributes through the quality of the intent they communicate.

    Marquet calls this moving authority to the information rather than moving information to authority.

    Where Pole A fits

    Pole A isn’t a mistake and fits a specific kind of situation.

    Genuine emergencies. The fire on the bridge. The outage at 3 a.m. that has the customer’s revenue line on the floor. The litigation hold that requires a single decision-maker with counsel within the hour. These are the conditions Marquet himself names as the exception: “emergency situations require snap orders.” Marquet himself reserved the authority to fire weapons because he didn’t want the conscience of taking another human being’s life on anyone’s mind but his.

    The vast majority of situations don’t require immediate decisions. Some do. A leader who can’t pull rank when the building is on fire is running a different failure mode.

    Competence not yet built. A new team may not have the competence Marquet names as the second pillar of leader-leader. “Control without competence is chaos.” When the technical understanding needed to make a decision well isn’t yet present in the team, the leader doesn’t get to delegate decisions that depend on that understanding. The work is to build the competence. Until then, the structure holding the work sits closer to Pole A than Pole B.

    Speed over quality. Some decisions need to be made now and don’t need to be optimal. The leader’s signature unblocks the work. “A little rudder far from the rocks is a lot better than a lot of rudder close to the rocks,” Marquet writes. But rocks exist, and sometimes a lot of rudder is the right move.

    Respect those situations. Pole A fits them.

    Much of the calendar of most senior engineering leaders, however, is filled with approval-type meetings that reflect none of the above. Those meetings come from a structure that hasn’t been redesigned for the work the team is actually doing.

    Where Pole B fits now

    Three changes have broadened Pole B’s fit.

    First, the information has dispersed. The senior engineering leader 15 years ago could plausibly hold the whole technical surface of a small company in their head. Today, in any organisation past about 30 engineers operating across cloud, data, and AI layers, nobody holds that surface any more. The information lives where the work happens. Authority that doesn’t live there is making decisions on stale, limited data.

    Second, the speed of consequence has risen. The interval between we made the decision and the consequence is visible in production has compressed from quarters to weeks and, in many cases, days or even hours. The Pole A bandwidth⁠—every non-trivial decision queueing at the leader’s desk⁠—was already slow when the consequence cycle was quarterly. With a daily consequence cycle, the leader can’t gate every decision and still ship.

    Third, AI agents in the loop change the scaling maths of authority. A leader running leader-follower with AI agents is asking the agents to wait for approval on actions the leader isn’t equipped to judge quickly enough. A leader running leader-leader with AI agents communicates purpose, constraints, and antigoals, then lets the agents⁠—and the humans working with them⁠—choose the actions. The right model supplies the competency, the harness context supplies the clarity.

    The quality of the leader’s intent becomes the variable that scales.

    An agent does twenty minutes of work, then waits a day for the approval, because the person who has to give it moved on to something else and won’t be back until tomorrow. The gate costs twice. It costs the wait, and it costs the leader the practice: the routine calls that built the judgement now go to the agent, and the leader who kept the approval kept only the accountability. Marquet’s word for the counter-move is connect, and the psychological-safety section below is what makes it possible.

    The doctrine is older

    Marquet didn’t invent Pole B. He made it work inside a US nuclear submarine.

    Military practice had already named the move under the heading of mission tactics: specify the mission and desired end-state, leave the action to the people closest to it, and accept that detailed plans go useless on contact while the intent survives.

    Commander’s intent—a crisp statement of the goal that never specifies so much detail that unpredictable events render it obsolete⁠—is the operational name for the language move that makes Pole B work.

    Credit, blame, and visibility all centralise under Pole A. If you are being promoted on the decisions that carry your name, the case for leader-follower writes itself. John Boyd used to put the choice to his protégés as a fork in the road: you can be somebody, taking the rank and the position and the rest of what the system hands out, or you can do something, which usually costs you the first. Pushing authority down is the second road. It gives away exactly the visibility the promotion runs on.

    McChrystal: the contemporary case at scale

    Stanley McChrystal gives us the contemporary case for leader-leader at scale in his account of the Joint Special Operations Task Force in Iraq from 2003 to 2008, written with Tantum Collins, David Silverman, and Chris Fussell in Team of Teams.

    The Task Force began as what McChrystal calls an awesome machine, a Taylorist clockwork that synchronised intelligence, operations, and analysis with assembly-line precision. It was superbly tuned for the world it had been built for. Al Qaeda in Iraq was defeating it through a network whose decentralised decisions ran at a tempo the Task Force couldn’t match.

    McChrystal’s two-part response carries the same teaching at the scale of a 7,000-person organisation in active combat.

    Shared consciousness: information transparency across teams, with the daily Operations and Intelligence brief eventually attended by all 7,000.

    Empowered execution: decision authority pushed to the level closest to the action, under a rule of thumb McChrystal describes this way: if something supports our effort⁠—as long as it is not immoral or illegal⁠—you could do it.

    This wasn’t the absence of leadership.

    The scaling problem he names⁠—too many decisions, too fast, with information distributed in ways the leader can’t hold⁠—is the one AI is now laying bare inside software organisations. Authority that lives with the leader scales linearly with the leader’s bandwidth. Authority that lives where the information lives scales with the quality of the intent the leader can communicate.

    Authority that lives with the leader scales with the leader’s bandwidth. Authority that lives where the information lives scales with the quality of the intent the leader can communicate.

    McChrystal had five years of command inside active combat to make the move. The CEO usually has a less unforgiving deadline, but the structural argument is the same. The move is the CEO’s to make, not the CTO’s.

    The cockpit-gradient story chapter 1 named via Korean Air 801 is the same move in another institution: lower the authority gradient so the first officer can call a missed approach without the captain’s permission.

    The cautionary case: Snook on the fallacy of social redundancy

    Delegation without the intent contract isn’t leader-leader. It is delegation theatre.

    Scott Snook names the failure mode directly. The fallacy of social redundancy is the assumption that adding more people to a decision automatically adds reliability. It doesn’t. If the structure doesn’t specify how the people share information, decide, and commit, adding them adds confusion, not safety.

    The leader who reads Turn the Ship Around! and concludes that the move is to stop giving orders is reading half the book. Marquet’s second pillar⁠—competence⁠—is the part that gets skipped. As Marquet puts it, control handed to people who lack the competence to use it well produces chaos, not leadership.

    The leader-leader move depends on technical and operational understanding being present in the people closest to the work. Where years of leader-follower structure have taught people not to think, the move must develop competence while the structure shifts. This isn’t a single decision made on a Tuesday morning. It is a programme of work that takes years.

    Snook’s caution is practical. A leader who delegates without the intent contract, competence-building, language discipline, and the psychological safety to exercise nominal authority produces a worse outcome than one who keeps the authority Pole B requires more discipline than Pole A.

    Marquet’s mechanisms⁠—21 specific moves

    Marquet’s first book names 21 specific mechanisms that made leader-leader work on Santa Fe: eight control mechanisms (the divest-and-distribute moves), five competence mechanisms (build the technical understanding authority depends on), seven clarity mechanisms (make sure people know what we are trying to do), and a closing 21st that stands outside the pillars.

    Two show how small changes rewrote authority.

    Find the genetic code for control and rewrite it. Marquet’s chiefs admitted they didn’t run the ship; he asked whether they wanted to. What they came back with was a signature. A leave chit is the form a sailor fills in to ask for time off, and Navy regulations said the XO signed every one. So a junior sailor’s request for Christmas leave collected seven signatures on the way up, four along his own enlisted chain and three more from the officers above it, then travelled back down again. Fourteen steps, on a form printed with five signature lines. Marquet found one of those requests sitting in an inbox while the sailor waited to learn whether he could book a flight home. The chiefs’ proposal was to strike XO from the regulation and write in COB, which moved the last signature out of the officer chain and into their own. One word. Marquet calls that word the genetic code.

    It didn’t stay one word. The chiefs could only own leave planning if they owned the watch bill, the roster that says who stands which watch and when. They could only own the watch bill if they owned qualification, the process that decides who is certified to stand each station at all. Managing leave, Marquet writes, was only the tip of the iceberg. The ship called the whole thing Chiefs in Charge.

    Use “I intend to…” to turn passive followers into active leaders. Marquet’s own account is that they eventually inverted the whole arrangement: instead of one captain issuing orders to 134 men, the boat became 135 independent, committed, engaged people all thinking about what needed doing and how to do it well.

    In Leadership Is Language, Marquet adds the andon cord: a preplanned pause signal, named in advance, that any team member can pull when something is off. It comes from Toyota’s assembly line, where pulling the cord lights the lantern that calls for help, the worker shifts from production to problem-solving, and the managerial discipline is to say thank you to the person who pulled it. Chapter 23 tells what happened when GM installed the cords without the conditions. Software organisations can use the same deliberate pause.

    Treat the 21 mechanisms as a curriculum: 21 operationally specified drills, each tied to a named situation Marquet walked through on Santa Fe over a single tour.

    The 21st is the one this book most needs and the summaries most often drop: Don’t Empower, Emancipate. Marquet’s verdict on the empowerment industry is that attempts to empower him always felt like manipulation, because the power was still being dispensed from above; emancipation recognises that the authority over your own work was never the captain’s to grant.

    If you are the CTO reading this book’s sidebars as a permission structure⁠—waiting to be authorised into leader-leader⁠—you are running the contradiction Marquet’s first chapter names. The mirror is something you offer the CEO. The authority over your own span you take.

    The enabling condition: psychological safety

    Marquet’s sixth play in Leadership Is Language is connect. Without it, the other five collapse back into their Industrial-Age forms.

    Connect requires what Amy Edmondson has documented for nearly three decades as psychological safety: the shared belief that the team is safe for interpersonal risk-taking.

    Without it, the I intend to statement doesn’t get made because the person who would make it is calculating the political cost of being wrong. Without it, the subordinate who sees something the captain has missed doesn’t say it because the cost of being seen to question is too high.

    The El Faro case study at the heart of Leadership Is Language is, in part, an Edmondson case: the third mate’s hesitant, self-diminishing, deferential, and nervous language pattern made it easy for the captain to reject the unwanted information, and the ship sank inside the next 12 hours.

    Without psychological safety, the leader running leader-leader is running leader-leader on paper and leader-follower in practice. The team’s behaviour is the data, not the leader’s stated principles.

    Edmondson’s toolkit for building the condition is specific enough to schedule, and it has three moves.

    Frame the work: say out loud, before the work starts, what kind of work this is⁠—how much uncertainty it carries, how interdependent it is, what failure rate is expected⁠—because a team told the work is novel treats error as data, and a team told nothing treats error as exposure.

    Invite participation: use situational humility (I haven’t run a migration like this before; what am I missing?), direct proactive inquiry at named people, and build structures⁠—the witness round of chapter 5, the dissent slot, the pre-mortem⁠—that make voice the default rather than an act of courage.

    Respond productively: thank the messenger before judging the message, treat the failure as the system’s output before anyone’s fault, and sanction clear violations openly, because safety isn’t the absence of standards.

    A mirror only does its work if the CEO can hear it. These three moves build the conditions in which hard information is received rather than rejected.

    Run them on your own team first. The team deserves it, and the CEO will learn the moves by watching you survive them.

    In decisions

    Last Tuesday shows what the move looks like. For a CTO or VPE, two surfaces matter: the CEO’s calendar and what that calendar enables or prevents one layer below, in the CTO’s calendar, the VPE’s calendar, and the senior leadership team’s behaviour.

    The CEO who runs leader-follower above the CTO often produces a CTO who runs leader-follower below: the structure inherits.

    Pole A leaders’ calendars fill with approval meetings. Pole B leaders’ calendars fill with conversations that begin I intend to… The conversation is shorter than the approval meeting, more frequent, and harder to schedule, because the team isn’t waiting for the leader. The leader is reachable rather than gating.

    Pole A leaders write task lists. Pole B leaders write a paragraph of purpose, a list of what must not happen, and a constraint envelope. The team writes the task list.

    Pole A leaders judge themselves by whether the decision was right. Pole B leaders judge themselves by the quality of the intent statement they communicated and the quality of the conversation when intent met reality.

    Pole A leaders are surprised when work stalls without them. Pole B leaders are surprised when work stalls with them, and treat the surprise as a signal.

    If you are reading this as the CTO or VPE, or you are the CTO who recognised yourself in this chapter’s opening paragraph, the structural fix is at the CEO layer.

    Installing leader-leader below a CEO who runs leader-follower above is the hardest version of the move, not an impossible one. Marquet’s first attempt, on Will Rogers, died exactly that death; on Santa Fe he made it stick by embedding mechanisms the structure couldn’t revert.

    What the CEO controls is whether your mechanisms compound or stay local.

    Put your own calendar from last week in front of your CEO, mark every approval meeting, and write one question under each: was this decision mine to make? Then do the same with the CEO’s calendar.

    The conversation isn’t about their leadership style. It is about the structure the CEO’s calendar is producing one layer below.

    McChrystal’s point holds: Shared consciousness and Empowered execution have to move together, and only the CEO can signal that Empowered execution now applies at the function-head layer. That signal is worth more than any number of leader-leader training programmes below it.

    The sentence your CEO can carry to the board: “Authority should live where the information lives.”

    The diagnostic move

    Three questions to work through this week. Solo. Not for sharing.

    Walk through your last week’s calendar. For each meeting where you exercised approval authority, ask: was the approval gating the work? Or was the work waiting for permission because the structure said permission was required? Count the second category. That number is the slack the structure has been hiding.

    Recall the last three decisions you made under time pressure. For each, ask: did the people closest to the work have the information needed to make this decision well? If yes, ask: why was the decision yours? If no, ask: what would it take to put the information where the decision belongs?

    Listen for the language your team uses when they bring you something. The phrases that map to Pole A are Request permission to…, I would like to…, What should I do about…, Do you think we should… The phrases that map to Pole B are I intend to…, I plan on…, I will… The ratio is the data. If you don’t hear any I intend to… in a week, the structure is likely running leader-follower regardless of what your stated principles say.

    In the language of chapter 5, these exercises compare espoused principles with observed behaviour.

    A team-based exercise

    Run one of Marquet’s mechanisms with your direct reports this week. For the CEO, direct reports are the senior leadership team: the CTO, the VPE if you have one, the CFO, the heads of function.

    The cleanest first move is the I intend to protocol on a single category of decision. Pick something⁠—a senior hire, a contract above your sign-off threshold, a strategic deal, a budget reallocation, an engineering deployment that has been routing through you⁠—and write the intent template with the team.

    The phrasing must come out as: I intend to [action], because [purpose], within [constraint], unless [antigoal triggered].

    Marquet’s own protocol is the first two slots plus an invited veto; the constraint and antigoal slots are my extension for AI-era delegation, where the agent can’t read the situation.

    For the next two weeks, decisions in that category arrive as I intend to statements. Your response is Very well, or⁠—if the intent is materially off⁠—a conversation about purpose, not a re-decision of action.

    If you are the CTO reading this and the CEO isn’t yet running the protocol with you, run it down: with your direct reports, on a category you have authority over. The structural argument still works one layer at a time, even when it can’t be made org-wide until the CEO joins in.

    Watch the team’s intent quality. It will be uneven at first, then improve fast if the protocol holds. Watch yourself for the urge to re-decide.

    Mine has a particular shape. What I am good at is translation, and its shadow is over-carrying: becoming the one who holds the integration for everyone. Every time I do that, the room gets one more decision it does not have to make and one less it knows how to make.

    The urge is information. It is the pattern that built leader-follower in you, still active. Notice it, name it for what it is, and let it pass without re-deciding.

    Going upstream

    In-text: Marquet, Turn the Ship Around!: the 21 mechanisms, the Santa Fe turnaround story, the operational curriculum. Marquet, Leadership Is Language: the language-as-lever distillation, El Faro as cautionary case, the six plays. Stanley McChrystal et al., Team of Teams: the contemporary case at scale, Shared consciousness, Empowered execution. Amy Edmondson, The Fearless Organization: psychological safety as the enabling condition.

    Also touched: Scott Snook, Friendly Fire: The fallacy of social redundancy and the cost of delegation without the intent contract.

    Go deeper: MCDP-1 Warfighting: the doctrinal source for mission tactics and commander’s intent. Robert Coram, Boyd: The Fighter Pilot Who Changed the Art of War: the Auftragstaktik lineage and Boyd’s to be or to do. Chip and Dan Heath, Made to Stick: commander’s intent in prose-accessible form for non-military readers.


    I work with engineering leaders on exactly this kind of paradigm work, the deeper the better. If it’s live for you, I’m happy to talk: schedule a 30-minute virtual coffee at hi.chrisgagne.com.

    Some book links here are Amazon affiliate links; if you buy through them I may earn a small commission, at no cost to you.


    On paper your values are Theory Y. The approvals and review gates your function runs on may not be. Your team reads the structure, not the values.

  • Batch and push vs single-piece and pull

    Batch and push vs single-piece and pull

    You set WIP limits and they’re being ignored. Your batches are getting larger. Your lead times are getting longer. The PMO says everything’s fine.


    Chapter 15 closed the safety cluster. This one returns to the work itself, where most planning still gets batch size wrong and pushes work onto teams beyond their sustainable capacity.

    🎧 Prefer to listen? This chapter is narrated in my own voice with ElevenLabs on Spotify (17 minutes).


    Chapter 4 asked whether the system was busy or moving. This axis asks how work is sized and routed, one level down in the same system.

    The two poles

    Pole A: batch and push. Aggregate work into larger units to amortise setup costs. Schedule and dispatch work to resources based on forecast. A central plan governs.

    Pole B: single-piece and pull. Reduce batch size. Resources pull work when they have capacity. Demand signal at the customer end governs upstream.

    Batch size and push-versus-pull are two separate levers. Batch size is how much work travels together, and it operates at three levels: the commitment horizon the board signs off, the size of a work item entering the team, and the size of a change reaching production. Push versus pull asks what triggers the next piece of work, a forecast schedule or free capacity downstream. WIP limits are the control that exposes both. The poles name the correlated ends of those levers rather than a single dial.

    Where Pole A is right

    When setup costs are genuinely fixed and high, demand is predictable, and inventory is cheap to hold: classical Economic Order Quantity (EOQ) logic in mature physical operations. The 1913 model assumes fixed setup and predictable demand, and prices the cost of holding inventory as an input. Raise that input and the model itself calls for a smaller batch. In the small slice of contemporary work that still meets them, the Pole A maths holds.

    A pharmaceutical manufacturing line with a multi-hour changeover and a regulated batch certification has fixed setup costs the operator can’t reduce. A semiconductor fab with a stable product mix has forecast accuracy a software team can’t match. The EOQ logic in those settings is sound maths. The planning most engineering organisations inherited was built for those settings, and the leaders who built it weren’t wrong about the world they built it for.

    Where Pole B is right

    Pole B is right wherever setup costs can themselves be reduced and the learning rate matters, where forecast accuracy is low, variation is high, and the holding cost of inventory is high. A deploy whose entire changeover is a CI run measured in minutes sits at the opposite end from the pharma line above. Software work with reducible setup cost, uncertain demand, or costly ageing inventory tends toward smaller batches.

    In decisions

    Pole A leaders schedule long roadmaps⁠—queues if we are willing to use a more truthful name⁠—and allocate capacity with long planning intervals. Pole B leaders ship continuously, set WIP limits, and let teams pull from a prioritised and pruned backlog.

    The batch-versus-pull axis is set by the planning the CEO runs with the board, and that is the conversation a CTO or VPE has to take upward. If the CEO commits to a quarterly or yearly roadmap with the board, the engineering organisation is pushed to batch work into large commitments no matter what the team-level discipline says.

    Put two measures in front of them: the cost of delay on one current initiative, in the CEO’s unit of account, and the change in cycle time the team achieved when WIP limits were last set and held. If you have never held one, the exercise at the end of this chapter produces that number in a sprint. Don’t turn this into an Agile lesson. Show the downstream cost: the planning they use with the board propagates into batch sizes the company and customer eventually pay for.

    Then ask for one change. What the board wants is confidence about risk and return; a 12-month feature list is a proxy for that, and a poor one, since I’ve never seen one delivered as written. OpenAI and Anthropic commit billions of dollars in compute years ahead and publish no dated feature roadmap. They commit capacity and direction, and refuse to commit the sequence. An incumbent with contracts and regulators can’t copy that outright, and doesn’t need to. Start with a count. How much of the roadmap is genuinely contracted, and how much is merely scheduled? In the organisations I see, the contracted share is small, and the rest had been treated as fixed anyway. Ask to take delivery rate and a record of what you redirected and why to the board instead.

    The sentence your CEO can carry to the board: “The planning calendar we run with the board becomes the queue our pending value stands in.”

    The planning calendar we run with the board becomes the queue our pending value stands in.

    The EOQ trap

    I suspect the Economic Order Quantity model is the least examined planning artefact in operations, and part of the reason is that almost nobody in software names it. The logic arrives second-hand, built into the planning calendar and the release train, and an assumption nobody names is one nobody audits. The 1913 model assumes fixed setup costs (which means setup reduction can’t change the result), accurate demand forecasts (which means uncertainty is small), and a holding cost that rises in a straight line with the batch (which means a big batch costs only proportionately more to hold). It then derives an optimal batch size from those assumptions.

    The maths is correct. Reinertsen’s own demolition runs through the cost legs: the “fixed” transaction cost isn’t fixed (Japanese manufacturers cut die changeovers from 24 hours to under ten minutes, using the single-minute exchange of dies (SMED) methods Shigeo Shingo pioneered), transaction costs actually grow with batch size, and holding costs grow faster than linearly.

    The trouble starts when one input no longer holds. If setup costs can be reduced⁠—through tooling, automation, practice⁠—the optimal batch size collapses. Low forecast accuracy does the same because the cost of being wrong about a large batch is much higher than about a small one. So does high inventory holding cost, especially when inventory has a short shelf life or hides quality problems. Setup cost is structural as much as technical. A method that convenes a two-day planning event for a whole programme every quarter carries a far higher transaction cost per batch than one that plans in a few hours each sprint. Hyper-specialisation compounds it. When every department is optimised on its own, finishing anything means crossing more of them, and each crossing adds coordination the batch has to carry. The dependencies raise the transaction cost, and the higher transaction cost argues for the bigger batch.

    Of the three, the setup-cost leg is moving fastest right now. AI-assisted engineering cuts the cost of producing and testing a change, which moves the optimum down the same way SMED did. My hunch is that it also moves the constraint, toward review, integration, and the demand governance deciding what gets built at all, none of which got faster at the same rate. If that holds, the WIP limit belongs where the work now waits: on review and integration, not on authoring. Don’t build faster than you can safely integrate, into production, into how the organisation works, and into your customers’ hands. If you can, spend the difference making integration faster and safer.

    The fact that the planning still defaults to legacy EOQ-shaped logic⁠—large batches, long forecasts, central planning⁠—is doctrinal lag. The maths says one thing; the org chart says another; the doctrine sided with the org chart. The critique predates Agile and DevOps: Goldratt’s The Goal ran it in narrative form in 1984, halving batch sizes against the economic-batch-quantity doctrine. Reinertsen supplies the general mathematics.

    Reinertsen on batch-size economics

    Donald Reinertsen’s The Principles of Product Development Flow gives the maths at industrial scale. It spells out how batch size, queues, and economic cost of delay lock together, with a precision the fragments of Lean and Agile most leaders inherit leave out.

    Batch size moves five variables at once: as batches shrink, cycle time, queue size, and risk fall while feedback rate and learning rise. The cost of reducing batch size⁠—usually some setup-cost-per-batch⁠—has to be weighed against the compounding benefit.

    Reinertsen’s running diagnostic comes from his own surveys of product developers, only 3% of whom had a formal transaction-cost-reduction programme. The cost-of-setup reduction is almost always achievable and almost always underinvested. It stays underinvested as a matter of inherited default, not because the economics favour the old batch size. Cheap switching is what makes adaptiveness affordable. The test is simple: after fresh evidence arrived, did the next item pulled actually change? A team can hold a hard WIP limit against a backlog nobody has touched in a year, and all the discipline buys is faster delivery of a stale plan.

    Reinertsen’s cost of delay, the economic argument chapter 4 made in full, applies directly to batch size. The visible cost (an engineer’s idle hour) wins planning conversations by default; the invisible cost, the delay a deep queue causes, never shows up until someone dollarises it.

    Anderson’s kanban

    David Anderson’s Kanban names the operational discipline. Anderson asks teams to make workflow visible, cap work in progress, measure and manage flow, state process policies aloud, and use explicit models to find where to improve. The board is the truth; the plan is the hypothesis.

    The failure I meet most often is treating kanban as a visualisation layer rather than a discipline. You adopt the board, the WIP limits get ignored, batches keep entering above the system’s capacity, and the conclusion is that kanban doesn’t fit your kind of work. When items keep entering above the cap, the board is serving as a status display rather than flow control. The WIP limit is the operative constraint. The board makes the constraint visible.

    The operating system

    Follow one item through. It is pulled when a slot opens rather than pushed when someone plans it, so it moves without waiting behind a batch. It ships small enough that the team learns something within days rather than at the end of a quarter. What they learn changes how the next one is pulled. Gene Kim’s DevOps Handbook names those three motions the Three Ways: Flow, Feedback, and Continual Learning. The first two you can build. The third is what the org design either permits or prevents, and it is what keeps the other two from eroding.

    In the State of DevOps data the fastest-deploying organisations, deploying on demand many times a day, also carry the lowest change-failure rates. The Pole A intuition that fast deployment is reckless is inverted by the measurement. Small change size is the likeliest mechanism⁠—less can go wrong in a five-line change than in a 500-line one⁠—though the surveys measure association rather than cause.

    A representative case

    The version I ran started somewhere less tidy than a planning problem. Demand was arriving well above delivery capacity. When we read queue length as time-to-clear at the observed rate, the answers came back in months and years. Nobody had been lying about the roadmap. The arithmetic had simply never been done out loud.

    Three changes.

    First, make the queue visible in time rather than in items. A count of stories is a number; most of a year to clear at the current rate is a decision.

    Second, put a hard edge on intake. Work went into two lanes with a stated order, and a displacement rule: when incoming unplanned work would push an iteration past the team’s raw velocity, enough planned work had to come out, with a comment on both stories recording what changed and why. The trade stopped being invisible.

    Third, set a work-in-process threshold and let it stop the line. More than 1.3 stories per builder, or more than three-quarters of an iteration’s worth in process, and intake froze. Both numbers were starting defaults, open to revision in retrospective, and they got argued about, which was the point: arguing about the threshold is arguing about capacity.

    The diagnostic move

    Three questions for last sprint’s allocation:

    • Which pole was I claiming? Did I describe the work as batched-and-pushed or single-piece-and-pulled?
    • Which pole would the actual batch sizes show? If I measured the average size of work items entering the system, would it look like Pole A or Pole B?
    • Which pole does this work actually require? EOQ wants a fixed setup cost, demand you can treat as known, and a holding cost that rises in a straight line. Miss any of those and the economic batch size drops. How far is a question for measurement, not doctrine.

    The exercise

    Run a WIP-limit experiment for one sprint. Set a hard team-wide WIP limit: say four items in progress total or half a sprint’s velocity, not per person. Pull from a ruthlessly pruned backlog. Measure lead time and cycle time before and after. The exercise produces three useful surprises. The team finds out which items have been passively blocked: sitting in progress because nobody had to kill them. The team feels how unfamiliar it is to pull rather than push. And the leader gets a clear lead-time signal that no roadmap-status conversation produces.

    A second variant for teams already running kanban: run a kata cycle on one constraint. Pick the stage where work piles up, name a target condition (how that stage should operate to produce a specific cycle time), identify the next obstacle, run a small experiment, learn, repeat weekly. Run the Toyota improvement kata as a four-week experiment.

    Going upstream

    Watch. Henrik Kniberg, High WIP, Context Switching vs One Piece Flow (Vimeo, 7 min). Don Reinertsen, GOTO 2012 Interview on flow (24 min). Companion short hook: Reinertsen, Economics of Batch Size and the “Father-Egg” Story (4 min).

    In-text: The two main sources named in the chapter. Don Reinertsen, The Principles of Product Development Flow (the Q-series, E-series, and B-series principles together, including batch-size economics and the EOQ demolition). David J. Anderson, Kanban (the WIP-limit discipline behind the exercise).

    Also touched: Eli Goldratt, The Goal, for the 1984 batch-halving antecedent to the EOQ critique (worked at full length in the local-optima instalment). Gene Kim et al., The DevOps Handbook (Chapter 1 on the Three Ways), the structure that lets Reinertsen and Anderson compose into one operating system.

    Go deeper: The lineage that names the same single-piece discipline at different layers, demoted here and carried elsewhere in the book. Mary and Tom Poppendieck, Lean Software Development (the seven principles for translating lean to software: cutting waste, building in learning, deferring commitment, shipping quickly, giving teams authority, designing quality in from the start, and optimising the whole rather than the parts) and The Lean Mindset, for the Toyota-Production-System-to-software translation. Jez Humble and David Farley, Continuous Delivery, for the deployment pipeline that makes large batches structurally expensive (developed in the utilisation-versus-flow instalment). Kent Beck, Extreme Programming Explained, for embrace change and the short-cycle cultural discipline. Mike Rother, Toyota Kata, for the coaching and improvement kata behind the exercise variant; the constraint-first selection above is the Goldratt overlay, since Rother starts from the pacemaker process, and he is precise that a number alone is a target, not a target condition. Gene Kim et al., The Phoenix Project, for the narrative form. Nicole Forsgren, Jez Humble, Gene Kim, Accelerate, the State of DevOps reports across four years, for the empirical correlation: high software-delivery performers are twice as likely as low performers to exceed profitability, productivity, and market-share goals, and the fastest-deploying organisations carry the lowest change-failure rates.


    I work with engineering leaders on exactly this kind of paradigm work, the deeper the better. If it’s live for you, I’m happy to talk: schedule a 30-minute virtual coffee at hi.chrisgagne.com.

    Some book links here are Amazon affiliate links; if you buy through them I may earn a small commission, at no cost to you.


    Small batches and pull only move work if someone holds the authority to decide what gets pulled, and in most structures that person is you. The queue forms on your calendar. Tuesday fills with approvals, the team stalls on any decision you haven’t signed, and the function gains no speed for all the oversight. You are the bottleneck. Chapter 17 names the default that put you there, leader-follower, and sets David Marquet’s leader-leader against it.

  • Safety as absence vs safety as presence

    Safety as absence vs safety as presence

    The incident counter your team reports up is at zero. What it can’t show you is how the work actually gets done safely and how to keep it that way.


    Chapter 14 walked through the anatomy of incidents. This chapter turns to what gets counted, because the safety metric you use decides what gets funded.

    🎧 Prefer to listen? This chapter is narrated in my own voice with ElevenLabs on Spotify (23 minutes).


    The deviation-versus-conditions question asks about a specific incident. Incident counts raise a system-level question: which view of safety drives the metrics and the investment? Both matter, and each demands its own focus.

    The two poles

    Pole A: safety as absence. Safety is the absence of events. The fewer accidents, incidents, and near-misses, the safer the system. Safety investment is insurance against accidents.

    Pole B: safety as presence. Safety is the presence of the ability to succeed under varying conditions. Safety investment is investment in adaptive capacity.

    Where Pole A is right

    Pole A fits tractable, complicated systems with stable specifications: the lost-time injury rate on a mature manufacturing line, the field-failure count for a long-standing physical product, or the audit-finding count on a stable regulated process. The absence metric tracks what’s happening because the variability is genuinely low.

    In those settings, Pole A is a strong metric because work-as-imagined and work-as-done are tightly aligned. The operations leader who uses it on a mature, stable manufacturing line is doing the right work. The same metric becomes structurally misleading in another domain.

    Where Pole B is right

    Pole B fits intractable, complex sociotechnical systems where work-as-done routinely adjusts to conditions work-as-imagined didn’t anticipate. A software platform under continuous deployment fits. So does a clinical service operating with staffing variability, or any system where the absence-of-events metric stops moving while everyone reports getting away with more than they used to.

    In decisions

    Pole A leaders count injuries. Pole B leaders ask, “How are we able to do this work well 999,999 times out of a million, and what do we need to preserve about that?”

    Safety-as-presence is a capital-allocation argument, and that is the version a CTO or VPE should carry up to the CEO. Put one picture in front of them: the four resilience potentials⁠—Respond (handle what’s arriving), Monitor (see what’s developing), Learn (extract the lessons), Anticipate (imagine what’s coming)⁠—mapped against where the engineering budget actually went last year. Hollnagel is careful not to prescribe a standard balance among the four; the right mix depends on the domain. I am one of four named inventors on PagerDuty’s operations-maturity patent, an instrument that scores incident practice from an organisation’s own operational events. My reading from that time is that engineering organisations are over-invested in Respond and under-invested in Anticipate (scenario planning, red-teaming, deliberate failure injection).

    In the engineering organisations I’ve seen, this split rarely came from a choice. Respond is urgent and visible; Anticipate is neither.

    The CEO-level move is to make Anticipate a named budget line, sized as a share of what you already spend on incident response rather than asked for as new money, so it survives the next planning cycle without requiring an incident to justify it. The exercises themselves don’t wait for the budget line; running them inside engineering is your span, and that part isn’t an ask: I intend to run a failure-injection exercise each quarter, because a green dashboard tells us nothing about the failures we haven’t met. The first is a tabletop, the second runs in staging, and we go near production only once we have shown we can stop one safely. Unless you see something I don’t, the first one runs this month. The sentence your CEO can carry to the board: “A quiet dashboard can’t tell skill from luck; we’re now funding the catching, not just the counting.”

    Safety-I and Safety-II

    Erik Hollnagel’s Safety-I and Safety-II names the framework. Safety-I⁠—Pole A⁠—studies what goes wrong. Its unit of analysis is the incident, and its improvement loop finds what caused the incident and seeks to prevent it from recurring. Its methods include fault-tree analysis, root-cause analysis, and near-miss reporting.

    Safety-II⁠—Pole B⁠—studies what goes right. Its unit of analysis is the everyday successful operation. Its improvement loop seeks to understand how things work when they work, and preserve the conditions that produce success. Its methods include appreciative inquiry, work-as-done observation, and resilience engineering.

    Hollnagel’s empirical claim is that Safety-I produces diminishing returns in complex sociotechnical systems. Most of the time, the system runs successfully despite gaps in work-as-imagined because people inside it make invisible adjustments. Studying only the failures misses what’s actually keeping the system safe. Safety-II keeps Safety-I and adds the missing half.

    Studying only the failures misses what’s actually keeping the system safe.

    The causality credo and the hypothesis of different causes

    Hollnagel names the unstated assumption under Safety-I: the causality credo. In Hollnagel’s words, it is the faith that “since all adverse outcomes have causes, and since all causes can be found and dealt with, it follows that all accidents can be prevented.” Closely related is what he calls the hypothesis of different causes: the belief that things going right and things going wrong have different causes.

    Resilience engineering rejects both. “Failures were the flip side of successes… things that go right and things that go wrong happen in basically the same way.”

    The same pipeline, the same time pressure, and the same workaround produce Tuesday’s save and Thursday’s outage; only the conditions differed. Successes also carry useful information. They are where the adaptive capacity lives.

    The Pole-B metric framework coheres only after leaders reject the causality credo. Pole B measures what works while it is working.

    The four resilience potentials

    Hollnagel proposes four potentials as necessary and, within the book’s argument, sufficient for resilient performance.

    Respond: knowing what to do: drawing on prepared actions, adapting how the system is currently running, or improvising new responses to whatever changes, disturbances, or opportunities arrive, whether routine or not. The capacity to handle what is happening now.

    Monitor: knowing what to look for: keeping watch on the things that do or could affect performance in the near term, both inside the system and out in the operating environment. The capacity to see what is developing.

    Learn: knowing what has happened: single-loop learning that draws lessons from specific experiences, and double-loop learning that revises goals or objectives. The capacity to extract lessons from both success and failure.

    Anticipate: knowing what to expect: looking ahead to possible disruptions, novel demands, fresh opportunities, or shifts in operating conditions further out in time. The capacity to imagine what is developing beyond what current monitoring can see.

    Together, the four form what Hollnagel calls the Resilience Assessment Grid. Organisations tailor its diagnostic and formative questions to their own work, apply them repeatedly, rate the answers on a Likert-type scale, and use a radar chart to track and manage change over time. This gives Pole B an operating structure. Safety is the presence of adaptive capacity; the grid names four potentials an organisation can assess and fund.

    Pole B leaders need to know which of the four potentials this organisation is strong in, which it is weak in, and how that investment changes over time. Most engineering organisations fund Respond and Learn—Respond because incidents demand it, Learn because retrospectives ritualise it⁠—while Monitor is patchy and Anticipate is starved by default.

    Cook’s points 16, 17, 18

    Richard Cook’s How Complex Systems Fail gives the three central points for this axis.

    Point 16: safety is something the whole system produces, not a trait of any one part. You can’t add up safe individual workers, safe individual procedures, safe individual tools, and arrive at a safe system. The safety emerges from how the parts interact under load.

    Point 17: People continuously create safety. Most of what keeps the system running is invisible adjustments by people inside it. The procedure said one thing; the operator did something subtly different that worked better in this specific situation; the system survived. This happens continuously. Pole A metrics don’t see it because the metric asks did an incident happen?, and no, the incident didn’t happen, because the operator adjusted. Pole B asks the adjustment to become visible.

    The procedure said one thing; the operator did something subtly different that worked better in this specific situation; the system survived.

    Point 18: keeping operations failure-free depends on people who have firsthand experience of how things fail. The skills the operator uses to keep the system running come from familiarity with how it can break. Reduce failure to zero and you also reduce the operator’s familiarity with the failure modes. The system becomes brittle in the long run because the people running it lose the experience they would need to recover when something does go wrong.

    The Pole A leader who succeeds completely produces a fragile organisation.

    Cook’s three points and Hollnagel’s four potentials teach the same thing at two scales. Cook names what safety is in a complex system; Hollnagel names how an organisation invests in it. The operator’s mid-incident adjustment is Cook’s Point 17 in the flesh, and Hollnagel’s grid is how an organisation buys more of it.

    Snook’s F-111 near miss

    Scott Snook’s Friendly Fire documents an instructive counter-case to the 1994 friendly-fire incident the previous chapter described. In September 1992, a flight of Air Force F-111s nearly engaged two Black Hawks from the same Eagle Flight detachment at Bashur, identifying them only on a second pass, in the Black Hawk pilot’s testimony: “At the last minute, the guy said those look like two Black Hawks.” Two years before the eventual shootdown, the same conditions nearly produced the same outcome.

    The 1992 near miss didn’t result in deaths. It also didn’t result in changes to the system. It surfaced only by chance, days later, over drinks at the Incirlik officers’ club, where the F-111 pilot told the Army aviators how lucky they were: his flight had been on the trigger and only then realised the targets were friendly. It was never reported up the chain. Nothing was done; nothing was learned.

    The near miss had the same structure as the eventual shootdown: no fighter-helicopter direct comms, fighters unaware of the helicopter mission, AWACS failed to relay. Two years later, the same conditions produced the accident.

    A Pole B organisation would have treated the near miss as critical data. A Pole A organisation, focussed on absence metrics, didn’t, because no accident had occurred.

    The near miss was the leading indicator. The accident report was the lagging one.

    Snook’s own reading is harsher still: even had the near miss been reported, he doubts it would have changed anything, because the organisational response to a report is to add rules, not to address the structural disconnects a report exposes. Reporting a near miss into a system that answers with procedure buries it a second time. Pole B demands investment in the cultural conditions where near misses get heard and the response is structural rather than procedural.

    Snook’s F-111 case shows what happens without a reporting culture and a just culture, the conditions under which people can surface what they saw without fear. Engineers don’t volunteer near misses to organisations that punish their existence.

    The AI overlay

    AI tools sharpen a Pole-A trap that was always there, and they do so in the way chapter 13 described through Bainbridge’s ironies. The automated dashboard reads incident counts as a primary signal. It can’t see the incidents operators prevented.

    As AI takes over more of the visible adjustment work, it strips out the slack the absence metric was living on. The metric never surfaced the invisible adjustments and human improvisations that kept the system running; now fewer operators remain to make them. Recovery gets harder.

    Organisations risk arriving at a state where the dashboards show everything is fine and the operators report being unable to recover when anything actually breaks. AI exposes how thoroughly Pole A had already hollowed out Cook Point 18: the absence metric had been spending down the operators’ familiarity with failure for years, and AI removes the last of the slack that hid the bill.

    Under AI, Pole B leaders actively study the augmented system’s failure modes while the dashboard is green. They use red-team exercises, deliberate failure injection, and post-incident reviews that ask what did the AI miss, and what would we miss if the AI weren’t here? Bainbridge adds another demand: preserve human access to the patterns automation has displaced, so human monitoring remains real. A green dashboard makes these exercises hard to fund. Pole B treats that lack of visible failure as the reason to run them.

    A representative case

    Suppose a large engineering organisation ends the year with no reportable incidents at all. The leadership team celebrates. A senior engineer privately raises a concern: the team has been working around several known issues for the entire year, and the lack of incidents has more to do with luck and operator skill than with system improvement. The leadership team has a green dashboard, a year of delivery wins, and a concern arriving without data behind it. They don’t act on it. The next year, two of the worked-around issues converge into a major outage that takes down the platform for nine hours.

    The post-incident review goes two directions at once. The Pole A direction looks for the proximate cause of the outage. The Pole B direction eventually asks what did the operators know in advance, and why didn’t the organisation hear it?

    Hollnagel’s Monitor potential had been missing from the organisation’s investment all along, whatever the dashboards suggested. The operators were monitoring; the organisation wasn’t. The Pole B answer reorganises how near misses get surfaced and resourced. The Pole A answer adds more monitoring of the wrong kind: more dashboards, more alerts. Both moves get made. Only one addresses the cultural condition that produced the year of unsurfaced workarounds.

    For engineering leaders, the year of zero incidents is the year to invest in Hollnagel’s Anticipate potential, provided you first run the test that says which kind of quiet year you had. Ask the operators what they worked around. A system whose operators report nothing may be genuinely well designed; a system whose operators have been working around known issues all year has been spending down its slack, and the incident count can’t tell the two apart. Snook, in the same chapter as the F-111 near miss: “Just because it isn’t broken, doesn’t necessarily mean that it isn’t breaking.” Use the quiet period to find out what the operators’ invisible adjustments are holding together. Pole A treats the quiet as success; Pole B treats it as the window during which Pole B investment is cheapest.

    The diagnostic move

    Three questions for last quarter’s safety review:

    • Which pole was I claiming? Did I treat safety as the absence of bad outcomes or as the presence of adaptive capacity?
    • Which pole would the actual investment show? Where did the engineering budget you steward actually go: into incident prevention or into resilience capacity? And across Hollnagel’s four potentials, which got the investment? Expect Respond and Learn to dominate, and check whether Monitor and Anticipate were funded at all. Is that distribution the one this system needs?
    • Which pole does the system actually require? If your system is complex and sociotechnical, expect Pole B to be the under-funded half, and let the audit in the next section say whether yours is.

    The three answers matter most when they diverge.

    The exercise

    Run a near-miss audit on the last quarter. Collect every near miss, every save, every moment the operators adjusted to keep the system running. Sort them by which of Hollnagel’s four potentials actually did the catching: Respond (the team handled it well when it arrived), Monitor (the team saw it developing), Learn (the team had been here before and knew what to do), Anticipate (the team had imagined this further out and was ready). When a save involves more than one, credit the earliest in the chain: if the team saw it developing, score it Monitor, whatever Respond did afterwards. Sorted that way, the distribution shows how early you catch things. The fat buckets are the potentials the organisation has built, whether or not it meant to. The thin ones are the potentials with the least practised cover. A fat Respond bucket is its own warning: the system is catching things at the last possible moment. A second variant for teams that have a strong retrospective practice already: run a Safety-II audit. What’s going right that the team doesn’t have visibility into? Where do operators tell you things work despite the procedure, not because of it? Those answers show where the Monitor potential needs investment, and they are the answers Pole A dashboards systematically miss.

    Going upstream

    In-text: Erik Hollnagel, Safety-I and Safety-II, and the operational companion Safety-II in Practice (Chapters 4 and 5 for the four resilience potentials and the Resilience Assessment Grid). Richard Cook, How Complex Systems Fail (points 16, 17, 18). Scott Snook, Friendly Fire, on the 1992 F-111 near miss (Chapter 7, the conclusion).

    Also touched: The reporting-culture and just-culture conditions that surface near misses are James Reason’s informed culture subcultures, worked at length in the deviation instalment (chapter 14).

    Go deeper: The sources under chapter 14⁠—Sidney Dekker, Safety Differently and John Allspaw, How Your Systems Keep Running Day After Day—both teach this dichotomy at the metrics-and-investment level too. The above-the-line / below-the-line frame is from Cook, Allspaw, Woods et al., STELLA Report (2017), unpacked in chapter 14. For pre-accident-investigation methodology: Todd Conklin, Pre-Accident Investigations.


    I work with engineering leaders on exactly this kind of paradigm work, the deeper the better. If it’s live for you, I’m happy to talk: schedule a 30-minute virtual coffee at hi.chrisgagne.com.

    Some book links here are Amazon affiliate links; if you buy through them I may earn a small commission, at no cost to you.


  • Deviation vs conditions

    Deviation vs conditions

    You’re looking for the deviator who failed. The deviator was doing what the system trained them to do.


    Bainbridge’s irony from chapter 13, that automation hollows out the operators it relies on, runs into incident response next, where blaming the operator leaves the conditions that produced the incident untouched.

    🎧 Prefer to listen? This chapter is narrated in my own voice with ElevenLabs on Spotify (30 minutes).


    The two poles

    Pole A: deviation. When work goes wrong, the explanation is a deviation from procedure, competence, or intent. The fix clarifies the standard or addresses the deviator. Human error is a sufficient diagnosis.

    Pole B: conditions. Real work continually adjusts to underspecified conditions. The gap between work-as-imagined and work-as-done is where both safety and risk live. Failure is the unexpected combination of normal variability. Human error is a symptom of that gap, not a diagnosis.

    Where Pole A is right

    Pole A fits tractable failures with a clean mechanism and locatable responsibility: a broken bearing, a software null-pointer exception, a structural calculation error, or an individual whose behaviour was genuinely reckless. Even that last item carries Dekker’s caveat. A proven bad apple is still a system responsibility: someone hired, credentialed, and scheduled them. Further, reckless is an infinitely negotiable label through which the deviation reflex re-enters, so treat it as where an investigation starts, not as its conclusion.

    Pole A is right when the failure is a discrete event in an otherwise stable system and responsibility can be named without distortion. Root-cause analysis works for that kind of failure. Many engineering-organisation failures aren’t that kind.

    Where Pole B is right

    Pole B fits any incident where failure came from the unexpected combination of normal variability.

    A hotfix I reviewed corrected one defect and introduced another: some transactions went through twice. The change passed every check the team had. The tester couldn’t run the production-like daily workflow in the test environment. The relevant edge case was absent from the manual test steps. There was no automated integration test for that path. The engineer making the change was carrying several high-risk fixes and preparing to hand work over, and the hand-off was rushed, with the important context living partly in private conversations.

    The monitoring detected the original failure mode. It didn’t detect a transaction that succeeded twice, so customers reported the result. The response then crossed engineering, service and product: every affected transaction had to be found, stopped where possible, put right and explained. The original defect still existed, because its fix had been reversed.

    Access, test design, workload, knowledge concentration, hand-off quality, monitoring and time pressure all contributed. Asking the engineer to be more careful would have left every one of those conditions intact.

    In decisions

    Pole A leaders trace the incident back to the decision that broke the standard, then reinforce the standard and coach the person who missed it. Pole B leaders go to the gemba—the actual place where the work happens⁠—and ask “what surprised you?” and “where did you have to improvise?” Then, they change the system people are operating within.

    Take your last meaningful incident and lay two analyses side by side. The first is your standard after-action review. The second report is what the on-call rotation knew beforehand and had no reliable channel to surface. That second report doesn’t exist until someone goes to the gemba and sources the knowledge, which takes effort and trust. The full protocol is at the end of the chapter.

    The gap between the two analyses is what Vaughan calls structural secrecy, a mechanism this chapter comes back to. Ask what this organisation has collectively stopped being able to see.

    The after-action review sits inside your own span, and that part doesn’t have to be an ask: I intend to redesign our reviews so the broader context is examined every time. Our current review isn’t going deeply enough, so unless you see something I don’t, the redesigned review starts with the next incident. The sentence your CEO can carry to the board: “What reaches us has been filtered; the deeper, unexamined context is where the next incident is hiding.”

    The Old View and the New View

    Sidney Dekker, in The Field Guide to Understanding ‘Human Error’, names two epistemic frames. The Old View holds that human error causes accidents. People who err are bad apples: remove them, retrain them, or stiffen the rules. The New View holds that human error follows from trouble inside the system. People who err are the symptom; the conditions that shaped their behaviour are what needs explaining.

    Dekker’s case, made across 30 years of safety-science work, is that the Old View’s empirical props⁠—accident-proneness statistics and Heinrich’s triangle⁠—fail under the data, while its remedies suppress the information safety management needs.

    Heinrich’s triangle holds that minor incidents and fatalities sit in fixed proportion, so a falling count of small events implies the big ones are under control. Incident dashboards run on that logic. Heinrich built the triangle from insurance and actuarial science: the ancestor of the modern incident dashboard is an underwriter’s loss table from nearly a hundred years ago. Dekker’s counter-example is Deepwater Horizon, where managers were celebrating six years of injury-free performance the day before the rig killed 11 people and caused the largest accidental marine oil spill in history. The record was real and accurately measured whether people were getting hurt on deck, but the well was a different risk class entirely. Construction shows the same split over decades: minor-injury frequency falling while the absolute count of fatalities and life-altering injuries holds steady. The ratios vary wildly between industries, jobs and places, and the safer a system gets, the less the small stuff predicts the large. Dekker’s charge is stronger than inaccuracy. The belief sedates: it “rocks you to sleep with the lullaby that the risk of major accidents or fatalities is under control as long as you don’t show minor injuries, events or incidents.” The quotation marks in Dekker’s book title compress the argument into typography: human error is a label someone applies, and labels aren’t categories of behaviour.

    The other prop is older, and it keeps coming back. In the mid-1920s the Boston public transport company found that 27% of its drivers caused 55% of its accidents, and psychologists in Britain and Germany independently proposed that some workers were simply accident-prone. Every engineering leader has seen the software version: one name recurs across the incident log, and the pattern looks like a person. The refutation is a denominator problem. The method needs every driver to face the same expected accident rate, and some drove the busy centre while others drove quiet suburbs at night. Dekker states it flatly: practitioners aren’t all exposed to the same kind and level of accident risk, which makes it impossible to compare their rates and conclude that personal characteristics explain the gap (p. 12). The engineer who appears in every incident review usually owns the gnarliest service. By the end of WWII the field had concluded that proneness was “much more a function of the tools and tasks he or she was given, and the situations they were put into” than of personal characteristics. The idea survives because blaming the worker leaves management “scot-free” of responsibility for design, selection and training.

    The Old View persists because it offers a fast, visible act at the moment the pressure is highest, and conditions-work offers nothing that week. Naming the deviator is easy and fast. Understanding, naming, and resolving the conditions is hard and slow.

    Naming the deviator is easy and fast. Understanding, naming, and resolving the conditions is hard and slow.

    Practical drift, at four levels

    Scott Snook’s Friendly Fire documented the 1994 friendly-fire incident in northern Iraq. Two American F-15 pilots, operating under Operation Provide Comfort, mistakenly engaged and destroyed two American Black Hawk helicopters. 26 people died. I think Snook’s book is the single most important book-length review of how organisational accidents actually happen.

    Snook refuses to pick one level like a Pole-A reading would. The pilots fired the missiles, so the pilots are the cause. Or the AWACS crew failed to challenge the engagement, so the AWACS crew is the cause. Or the rules of engagement were ambiguous, so the rule-makers are the cause. Each was true at its level, but none is the whole story.

    Snook’s four-level causal map names the structure: individual (what the pilots saw and did), group (what the AWACS crew did), organisational (how the no-fly zone command structured decision rights), and cross-level (how the layers interacted under conditions of practical drift). Each layer has its own causal story, none sufficient on its own to explain the accident.

    Picking one level is, in Snook’s framing, how resolution gets felt without being reached. The simple story picks one level and calls it the cause; the multi-causal story holds all four at once.

    Snook named the pattern of small-step accommodation across these levels practical drift: “the slow steady uncoupling of local practice from written procedure”. The 2×2 he drew has two axes: is the action rule-based or task-based, and is the system loosely or tightly coupled. Four quadrants follow: Designed, Engineered, Applied, Failed. On Snook’s account, most organisations live in Applied, where local task logic and loose coupling keep everything working.

    Normal Behavior, Abnormal Outcome is his heading for what the move from Applied to Failed produces when coupling suddenly tightens. Everyone is doing what local practice says. The aggregate produces an outcome no one actively chose. The matrix is a cycle: the organisation that has just been to Failed writes tighter rules, which is the move back to Designed, and the drift begins again from there.

    Snook’s verdict is that the shootdown was, in the strict sense, normal. Normal people behaved in normal ways inside normal organisations, exactly as the theory would predict given the circumstances each level faced.

    The Pole A response to practical drift looks for the moment someone broke a rule. There was no single decisive moment. Pole B asks one pair of questions at all four levels: what did work-as-imagined say should happen, and what had work-as-done drifted into?

    Structural secrecy: what makes drift invisible

    Snook’s practical drift named the pattern. Diane Vaughan’s The Challenger Launch Decision names what makes the pattern hard to recognise as a signal from inside the organisation.

    Structural secrecy is Vaughan’s name for the way an organisation’s own patterns of information, structure, processes and regulatory relations undercut its attempt to know and interpret its own situation. That runs at every level, not only the top. It is the ordinary functioning of a multi-level organisation, without concealment by intent.

    That hotfix ran on exactly that. The context needed to catch the edge case lived in private conversations during a rushed hand-off, so it was in the organisation and not in the review. The monitoring watched for failures, and a transaction that succeeds twice is still a success, so the system had no way to see what it was doing. Nobody concealed anything. Each representation showed a true slice, and the slices didn’t add up to the situation.

    Structural secrecy is the one of Vaughan’s Challenger forces a senior leader most needs. In Vaughan’s own case the channels were open and the top knew what the work group knew; the filter was as much in the shared construction of risk as in transmission.

    Dekker carries Vaughan’s concept into the safety bureaucracy itself: the cultural, organisational, physical and psychological separation between operations on one side and safety regulators, departments and bureaucracies on the other, a gap that widens as the safety function grows its own remit and budget. Clarke and Perrow’s 1996 fantasy documents sit in that gap: safety plans that bear no relation to actual work, written to persuade regulators and boards that the hazard is handled. They are the paperwork that makes an organisation feel like it can see.

    Structural secrecy and practical drift feed on each other. Not knowing what other units do lets work groups drift into locally practical arrangements, and the distance that opens between groups deepens the secrecy in turn.

    Cook’s 18 points

    I suspect Richard Cook’s How Complex Systems Fail (a short treatise first published in 1998) is the densest one-document statement of the New View. It contains 18 numbered claims. Three matter most for this axis; the point titles below are my paraphrases, not Cook’s.

    Point 7: There is no isolated root cause for failure in a complex system. Failure emerges from multiple latent conditions interacting under specific circumstances. The search for a single root cause is, in this framework, a categorical mistake. The system produces failure the way a wet floor produces a slip. The floor alone is harmless. It takes the wet floor, the smooth shoes, the box someone is carrying that they can’t see past, and the hurry, all present at once, and the slip is what that combination does.

    Point 8: Hindsight biases post-accident assessments. After the fact, the path to failure looks obvious. Why didn’t they see it? They couldn’t see it from inside the situation because the path became obvious only in retrospect. Further, what I’ve seen in incident records isn’t that late reviews get vaguer but that they mostly don’t happen: the document exists as a title wrapped around an unfilled template. The hotfix above was itself the fix for an earlier incident, deployed earlier that same week. Whatever review that earlier one was owed, the next incident arrived first.

    Point 15: Post-accident remedies often increase coupling and complexity. The intuitive Pole A response⁠—add a procedure, add an approval gate, add a monitoring system⁠—frequently makes the system more tightly coupled and harder to operate safely. The remedy can make the next incident more likely. Cook’s claim is hard to act on because the alternative⁠—sit with the incident, study the conditions, do less⁠—violates Pole A’s instinct. “Don’t just do something, stand there (in the Ohno circle)!”

    Point 15 names a design choice as well. You can try to make the system unable to fail by listing every way it could fail and preventing each one, or you can design it to contain failure when it arrives, because no list is complete in a self-organising system. Alicia Juarrero’s constraint architecture, developed at length in the local-optima chapter, is the philosophical statement of the same choice. Punishing the operator leaves that architecture untouched, which leaves the next incident pre-loaded.

    Behind Human Error

    Most incidents carry two stories. In one review I read, the first story arrived in the middle of the night, while the outage was still running: a senior leader confirmed human error as the root cause and instructed the team accordingly. It was the fast answer, available at the moment the pressure was highest, and it was wrong. The considered analysis, written later from the same facts, put the cause elsewhere: no change control for the production system, no runbook, no completion checklist, no post-task validation. Nothing about the timeline changed. The explanation did. The second story is the one underneath it. The procedure didn’t fit the conditions the operator was facing; the operator improvised; the improvisation worked most of the time… until it didn’t.

    Keep working until you have the second story. Sharp end—the operators in the moment⁠—and blunt end—the organisational conditions shaping what the sharp end can do⁠—must both be in the analysis. Punishing the sharp end while leaving the blunt end unchanged is the failure I see most often in the review records I’ve read.

    Punishing the sharp end while leaving the blunt end unchanged is the failure I see most often in the review records I’ve read.

    The boundary the organisation draws between operator and system is political, and that choice determines who gets blamed. Pole A draws the line so the operator stands alone. Pole B includes the system in the analysis.

    Work-as-imagined and work-as-done

    Above the line is what the organisation describes about how the work gets done: org charts, policies, procedures, dashboards. Below the line is what actually happens when the work is being done: the improvisations, the workarounds, the implicit knowledge, the social adjustments. The gap between the two is where safety lives, and where risk lives. (The phrase is adapted from the STELLA Report’s line of representation, which draws its line in a different place: the people and their mental models above it, the code and hardware below.)

    The pair of terms is Erik Hollnagel’s. After-action reviews that stay with work-as-imagined produce policy clarifications. Reviews that reach work-as-done produce systemic learning. Spend enough time in the work for work-as-done to become legible.

    Vaughan’s structural secrecy operates here. The organisation’s above-the-line representation is the documented work. The below-the-line reality is the actual work. Structural secrecy lets both versions coexist without anyone recognising the gap as a signal.

    Reason’s informed culture

    James Reason names the cultural toolkit Pole B requires. An informed culture—one in which leadership knows what is actually happening at the sharp end⁠—rests on four interlocking subcultures: a reporting culture where people freely surface errors and near-misses without fear of inappropriate blame; a just culture that distinguishes honest error from the small minority of behaviours that warrant sanction so the reporting culture can survive; a flexible culture that can move authority to the sharp end during high-tempo operations; and a learning culture that turns lessons into reconfigured assumptions and action.

    Reason’s just culture carries a caveat. It assumes the line between honest error and sanctionable behaviour can be drawn clearly and consistently. Dekker’s Just Culture argues there is no line, only people with the power to draw it: the same negotiability that hangs over reckless earlier in this chapter. Treat the distinction as something your organisation negotiates in the open. The taxonomy won’t hand it to you.

    Some of the engineering organisations I’ve worked with claim what Ron Westrum called a generative culture and operate, by Reason’s four-subculture checklist, as bureaucratic ones. Reason’s framework tells leaders which subcultures must be present before that claim is empirically true.

    Reason’s separate GEMS taxonomy supports a more precise discipline. The same outcome⁠—the operator did the wrong thing—has at least three different mechanisms underneath, and each needs a different response. A skill-based slip is the config pushed to the wrong environment because the two consoles look identical: the response is a guardrail in the interface, not more concentration. A rule-based mistake is the runbook written for the old topology, so the right rule fired on the wrong system: the response is fixing the runbook and teaching when it applies. A knowledge-based mistake is the novel cascade with no runbook at all, where someone improvises a plan that turns out wrong: the response is real support for sense-making in territory nobody has mapped.

    The Swiss cheese model is itself a simple story

    Deviation and conditions are also two ways of telling the story of a failure. The deviation account is a simple story: one cause, one villain, one fix. The conditions account is multi-causal: several causes operating at once, no single villain, a fix that touches the system. The narrative form determines the intervention that follows, which is Jennifer Garvey Berger’s simple-stories mindtrap operating on an incident review.

    James Reason gave the field what became known as the Swiss cheese model: defensive layers with holes that vary over time, and accidents happening when the holes align. The metaphor crossed every tradition in the field, from aviation to healthcare to software. Unsafe acts by front-line workers still do real work inside Reason’s model, and the phrase is Heinrich’s. Dekker dates the thinking underneath the diagram as “getting on in age⁠—soon a century” (p. 123).

    The metaphor itself tells a simple story about multi-causal failure. Linear slices. Discrete holes. Causation moves left-to-right through the diagram. The model meant to teach multi-causality is the most Pole-A-shaped way to talk about it.

    Even the field that argues hardest for multi-causal stories reaches for a linear metaphor when it has to communicate. The Swiss cheese model travels because it is simple. Its simplicity makes it portable and limits what it can teach.

    The Pole B move is to use the model as an entry point and refuse to stop there. When the after-action review reaches for the Swiss cheese diagram, ask which non-linear interactions, dynamic feedback loops, and structural conditions the linear diagram couldn’t depict. Carry Garvey Berger’s own summary into that room: “in a complex world a simple story is just about always wrong, and will just about always lead us to an emaciated, impoverished set of choices” (p. 39). Her habit is the discipline that room needs: notice your story, then create another, then another.

    A representative case

    Picture a platform engineering team a few months into a run of deployment outages. The Pole A response from leadership is clear: tighter change-control, more approval gates, after-action reviews focused on what each on-call engineer did wrong. The outages continue. The engineers are exhausted and starting to leave. The ones who were on call when it broke carry it hardest: Dekker calls practitioners harmed by an event they were caught up in second victims, and support for them is an obligation.

    When the Pole B move lands, a senior engineer is given a week to do nothing but interview the on-call rotation and the platform team about what the work actually looks like below the line. What the interviews surface is consistent. The change-control system has become so onerous that engineers are batching small changes into larger ones: fewer change reviews, but each change touches more surface. The approval gates have introduced a delay that pushes deployments into peak-traffic windows. The after-action reviews have become defensive theatre that makes engineers reluctant to surface near-misses. Reason’s reporting subculture, which the leadership team thought it had, is structurally absent below the line.

    Each Pole A remedy has a mechanism running the wrong way: batching enlarges the blast radius, gate delay pushes deploys into peak traffic, and defensive review suppresses the near-miss reports leadership needs. That is Cook’s Point 15 in operation. The Pole B response⁠—reduce change-control friction, redesign after-action review to surface conditions rather than assign blame, give the on-call rotation slack, build a real reporting subculture by making the engineers who surface near-misses visibly protected⁠—is the one to test against the outage rate. Nothing in the interviews pointed at the engineers; everything pointed at the conditions they were working in. And blame-free still carries accountability: in Dekker’s terms, the account stops being something the operator settles and becomes something the operator tells. Engineers in a blameless review are, in John Allspaw’s phrase, very much on the hook for helping the organisation become safer.

    The diagnostic move

    Three questions for last Tuesday’s incident:

    • Which pole was I claiming? Did I treat the incident as a human deviation or systemic condition?
    • Which pole would my actual response show? Did I ask who failed? or What surprised you? Where did you have to improvise?
    • Which pole does this incident actually require? If the failure was the unexpected combination of normal variability, Pole A is structurally inadequate. If it was a single clean cause with a locatable mechanism, Pole A is apt.

    The exercise

    Run a second-story protocol on the team’s last meaningful incident, with one constraint: spend at least half a day below the line at the gemba before the after-action review gets written up. Talk to the operators. Ask the two questions Pole B leans on: what surprised you? and where did you have to improvise?

    Map what you find at four levels⁠—individual, group, organisational, cross-level⁠—without forcing a single root cause. Compare it with whatever first story the organisation had already produced. The exercise lives in noticing the gap.

    A second variant for organisations with mature reviews is to place the incident on Snook’s matrix. Was the action rule-based or task-based, and was the system loosely or tightly coupled? That lands you in Designed, Engineered, Applied, or Failed. Snook’s theory predicts you will land in Applied without having seen it. Naming the quadrant tells you which way the drift is running, and whether the remedy you are about to write starts the next lap.

    Going upstream

    In-text: the primary sources named in the body. For the Old View / New View frame, and for the dismantling of Heinrich’s triangle and the accident-proneness data: Sidney Dekker, The Field Guide to Understanding ‘Human Error’. For the multi-level anatomy and practical drift: Scott Snook, Friendly Fire. For the upstream mechanism, structural secrecy (with the production of culture and the culture of production, the other two forces in her Challenger analysis): Diane Vaughan, The Challenger Launch Decision.

    Also touched: For the 18-point statement of the New View, including the no-root-cause and increased-coupling points: Richard Cook, How Complex Systems Fail. For the cultural toolkit⁠—reporting, just, flexible, learning subcultures⁠—and for the Swiss cheese model itself: James Reason, Managing the Risks of Organizational Accidents. On blameless review and the engineer as the person on the hook: John Allspaw, Blameless PostMortems and a Just Culture. On the negotiability of the just-culture line: Sidney Dekker, Just Culture. For the simple-stories habit of carrying two more stories: Jennifer Garvey Berger, Unlocking Leadership Mindtraps.

    Go deeper: the convergence pile-ups and further reading, with no in-body anchor. The no-single-root-cause claim arrives from several traditions beyond Cook and Snook⁠—Juarrero, Woods et al., and Hollnagel⁠—each naming it differently. For the second story, sharp-end / blunt-end, Westrum’s pathological / bureaucratic / generative typology of organisational cultures, and interface as line of demarcation for blame (after Kelly-Bootle, 1995): Woods, Dekker, Cook, Johannesen, Sarter, Behind Human Error, second edition. For the line of representation this chapter adapts: Cook, Allspaw, Woods et al., STELLA Report. For the GEMS / skill-rule-knowledge error taxonomy: James Reason, Human Error, and Jens Rasmussen’s SRK framework, carried in Behind Human Error. For the philosophical foundation, enabling constraints and the fail-safe / safe-fail choice: Alicia Juarrero, Dynamics in Action, developed at length in the local-optima and leader-leader chapters. For negative capability, the capacity to remain in uncertainty without an irritable reaching after fact and reason: John Keats’s December 1817 letter to his brothers. Watch. Sidney Dekker, Safety Differently (2018 lecture, 33 min); software-native companion, John Allspaw, How Your Systems Keep Running Day After Day (DevOps Enterprise Summit 2017, 33 min). Read. Sidney Dekker, Drift Into Failure; Erik Hollnagel, Safety-I and Safety-II; Todd Conklin, Pre-Accident Investigations; James C. Scott, Seeing Like a State; W. Edwards Deming, The New Economics, on variation.


    I work with engineering leaders on exactly this kind of paradigm work, the deeper the better. If it’s live for you, I’m happy to talk: schedule a 30-minute virtual coffee at hi.chrisgagne.com.

    Some book links here are Amazon affiliate links; if you buy through them I may earn a small commission, at no cost to you.


    That was the incident-level call. The system-level call asks what the organisation counts as safety. The incident counter your team reports up reads zero, and you take that for safety, when all it records is the absence of whatever got counted. It stays silent on how the work goes right on the ordinary days, which is what keeps you safe. Chapter 15 sets safety as absence against safety as presence, and reaches for what a leader would measure in its place.

  • Your AI transformation will fail the way your Agile one did

    Your AI transformation will fail the way your Agile one did

    Nearly nine in ten change programmes fall short of what they set out to do. Your AI transformation is being set up to join them, and the reason has almost nothing to do with the technology.


    A bonus edition of Come Prepared to Die, outside the numbered chapters.


    Look squarely at the last big change programme you ran or sat inside. Agile, most likely, or the DevOps push, or the one with the framework and the certified coaches and the reorganised standups. My read, over two decades in and around this work, is that almost all of them fell short of what their framers intended when they wrote the Manifesto, and that most of us who led them know it privately even where we defended them in public. Bain looked at 250 change programmes and found 12% achieved what they set out to do; the rest split between failing outright and settling for a shortfall they learned to call success.

    Agile didn’t fail because Agile was wrong. Your AI transformation is being set up to fail for the same reason.

    Agile was never a methodology, it was an organisational capability

    You don’t get agility by running the ceremonies any more than you get a butterfly by pinning wings to a caterpillar. Jay Galbraith gave us the picture decades ago, and it still holds. He called it the Star Model and laid an organisation’s design out as five points pulling on one another: strategy, structure, rewards, processes and people. An organisation’s behaviour emerges from how all five are arranged together, not from any single one of them. Further, whether that behaviour is worth anything turns on whether it fits the situation the organisation is actually in. An organisation design tuned for a stable world produces orderly behaviour that is exactly wrong for a volatile one.

    Software development—especially now that AIs are involved—is volatile, uncertain, complex and ambiguous, which is the VUCA the war colleges coined for exactly this kind of ground. The behaviours that pay there are adaptive. It’s true that some of the work is merely complicated rather than genuinely complex, and there an orderly design still earns its keep; but betting on the organisation staying in calm, predictable water is a bold call. Better to assume complexity and downgrade to orderly than to assume orderly and get thrashed about. The Org Topologies work Craig Larman does with Alexey Krivitsky and Roland Flemm builds straight onto the Star Model, and it treats agility as a property of the whole design, not a set of practices a team can adopt.

    Before the design argument, the money one. AI cuts task time. On its own it doesn’t move revenue, margin or cash. The gain converts only if the freed capacity gets redeployed to something the business sells, and that decision sits across function boundaries, which puts it with whoever holds P&L authority over the whole workflow rather than with the department that bought the tool. Where nobody holds that, the capacity turns into idle time or more work-in-progress, and the company pays for licences and a change programme while the benefit leaks between functions.

    The first wall: what you’re allowed to touch

    The thing an Agile transformation is trying to move lives in the whole organisation, and the effort is scoped to one corner. The people brought in to run the change, the coaches and the transformation office, generally only touch local processes inside product and engineering. The wall is porous, not solid: a CTO or CPO with real authority inside their own department can move some local structure, some of the people, even some of the local rewards, and a change effort they back reaches that far with them. What none of them reaches is the design of the organisation around the department, the reward system that sits above it, or the behaviour of stakeholders who never report in. For those, the change effort holds neither the remit nor the standing.

    And where it does push, the system pushes back. Larman put this in his Laws of Organisational Behaviour years ago: organisations are implicitly optimised to protect the power and positions of middle managers and specialists, so any change initiative gets reduced to redefining the new terminology to mean roughly the old status quo. So the boundary works as a machine as much as a fence. Push a change through it and the change comes back out wearing the new words and the old shape. Nobody built it on purpose. It defends itself anyway, and very well.

    Larman and Bas Vodde have a line for what the wall costs you: be agile, not do agile. Doing agile is the local practice a team can run inside its own boundary, the part an outside change effort is allowed to install. Being agile is a property of the whole star, and reaching it means changing the points that effort was never given the authority to touch.

    The second wall: how deep you’re allowed to go

    Even inside what it can touch, the effort is held at the surface. Peter Block, in Flawless Consulting, separates the content level of a consultant’s work, the frameworks and recommendations a manager can absorb without changing how they see themselves, from the affective level, where trust, power and identity actually live. Real change doesn’t happen unless the affective level shifts as well. Most consultants won’t go there, Block says, because it makes them vulnerable, and they collude with the client in pretending the organisation is a rational machine rather than a political one. Jerry Weinberg drew the same wall as a number in The Secrets of Consulting: never promise more than about 10% improvement, because 10% is the most you can deliver without threatening the client’s paradigm. The ceiling is set exactly where the discomfort would start.

    Consultants and coaches get blurred together here, and they are trained for different halves of the problem. Consultants work at the content level. They bring the diagnosis, the framework, the recommendation, and most of them have neither the training nor the invitation to work at the affective level at all. Coaches are built the other way round. They are trained for the affective level, for what a person believes and fears and protects, and most of them don’t carry the technical content of how a software organisation is actually put together. The rare practitioner holds both, and getting there takes two separate educations, run by different institutions and certified by different bodies. Nothing in either pipeline produces the other. I took both. It bought me no authority at all: neither trade, alone or combined, was ever handed the power to change the design.

    Above that wall sit the leader’s own paradigm, the mostly-unexamined model of what an organisation is and how it should work; their read of the context they are actually in, which may be years out of date; and the goals they are optimising for, which may not be the ones the situation rewards. Those aren’t the only inputs to a strategy, and the arrow doesn’t run only downward. The design already in place decides which strategies are thinkable at all, and a company whose rewards and reporting lines were built for the old model will keep choosing strategies that fit it. These three matter because of where they sit: above the wall, on the leader’s own ground, and no content-level intervention reaches any of them. A consultant can redesign a standup. A coach can hold the room while a CEO looks at what they believe an organisation is for. Neither one can do the looking for them.

    The two walls together box the effort into the smallest, lowest part of the star: local in scope, shallow in depth. The whole-system property it is chasing is a feature of the entire design and of the identity above it, not of any single point you are allowed to touch. More consulting, better consulting, another framework, and you are still inside the box.

    The same trap, one rung up

    I’m not the only one saying this. Stefan Wolpers, writing on the Scrum.org blog in January 2026, calls it the Agile-AI isomorphism: organisations that installed Scrum ceremonies without changing structure, culture and governance failed at Agile, and organisations installing AI tools without changing those same conditions risk failing at AI. The predictor of success, he argues, is whether the organisation genuinely changed last time or merely bought the process.

    So what has the organisational response to AI been so far? IBM surveyed 2,000 CEOs early in 2026 and reported that 76% of their organisations now have a Chief AI Officer, up from 26% a year earlier. Read quickly, that looks like the authority wall coming down at last. A C-suite owner, finally.

    I read it the other way, and I think the number itself gives me the grounds to. Mintzberg, citing Rumelt’s survey of the Fortune 500, records the last structural change of that size: divisionalisation went from a fifth of those firms in 1949 to three-quarters by 1969, and he files even that under fashion rather than fit. Twenty years, not twelve months. What can happen that fast is a title. A CTO takes the AI hat, a VP gets re-lettered, a CEO ticks “yes” on a survey run by a vendor that sells AI transformation. The evidence is in the same study. IBM reports 76% with a Chief AI Officer while only about a quarter of their people use AI regularly at work. The title has run a full lap ahead of the work.

    And the work, where it runs at all, is mostly not paying its way. MIT’s NANDA initiative reviewed about 300 publicly disclosed enterprise AI initiatives through 2025, interviewed 52 organisations and surveyed 153 leaders, and found 95% showed no measurable impact on profit and loss. Take that number at the weight it can bear: the work is preliminary, it hasn’t been peer-reviewed, the sample isn’t random, and a short measurement window will under-count anything still maturing. The models were rarely the binding constraint. What failed was the integration into how the business actually works.

    NANDA didn’t test authority, mandate, or the leader’s paradigm, and I’m not going to borrow certainty from their result. That reading is mine, so I checked it. A title isn’t evidence of authority unless the charter carries decision rights over structure, rewards and cross-functional organisation, so that is what I coded across 50 AI-transformation roles at 49 organisations. I registered a prediction before I looked and got it wrong: I expected at least 70% scoped to the AI function alone, and it came in at 64%. The other half held. Two roles in 50 carried written authority over structure, incentives, cross-functional design or P&L. Of the 13 reporting to a CEO or board, none did. The altitude was real; the remit wasn’t. These are charters as written rather than authority as exercised. The rubric, the 50 rows and every coding call I had to defend are written up separately.

    I wanted to run the same coding against the Agile era. If both eras chartered the work below the level where the design gets changed, the parallel stops being an analogy and becomes a measurement. I couldn’t do it. The coding needs the role’s reporting line, and for the Agile years that field has mostly gone from the public record. So the Agile half of this stays an argument rather than a finding.

    That is Larman’s Law arriving on schedule, one rung higher than the Agile PMO and wearing a better suit. What the 76% measures is how fast an organisation can announce it is Agile AI-native. Copilot is Jira with a larger budget and the same blind spot.

    Worse than Agile in two ways

    The 10X Org authors have a name I keep borrowing: the Ferrari Effect. Buy a fleet of Ferraris and put them on a one-lane road behind the trucks, and all you get is faster cars in the same traffic. AI dropped into a misaligned structure does exactly this. It amplifies whatever the organisation was already optimising for, so if the design was tuned for utilisation and local efficiency instead of flow and adaptiveness, AI makes you worse at the wrong thing, faster. Nothing an organisation does converts into performance except through its fit with the situation, and behaviour and culture both pay that same toll. A faster car doesn’t improve the fit. Fix the road first.

    The first way is that AI doesn’t stay neutral while you misuse it. Trained on the internet’s management writing, it hands back the fashionable answer more readily than the fitting one. Researchers writing in the Harvard Business Review in early 2026 ran thousands of simulated strategy decisions through leading models and found they reliably reached for whatever matched current management language; they named the output trendslop. Agile at least failed in silence. AI fails while affirming you, in fluent and confident prose, that you are doing the right thing. The one tool you would want to expose your paradigm is the tool most likely to sell it back to you.

    The second way is that AI widens the gap between a startup and an incumbent further than Agile ever could. A startup carries none of the incumbent’s inertia. It has no legacy design to defend, and at that size the founder simply is the structure, so no wall stands between the engineers and the design of the company. There is no gap between the paradigm and the design, and nothing to revert to. Every incumbent was a startup once. The org design grew organically after that, along the old hyper-specialisation lines, optimising for efficiency and utilisation rather than for delivering value, and each layer of that growth is now something with a constituency to defend it. The startup begins in the configuration you keep reverting away from.

    In the Agile era that gap had a ceiling. A well-run startup could out-ship an incumbent and still lose, because reaching scale took people and the incumbent had them. Headcount was the incumbent’s answer to speed, and it was a good one. My read is that AI moves that exchange rate. When leverage per person rises far enough, a small organisation whose design fits its situation can reach a size that used to need a large one, and headcount stops being the answer it was. The compensating strength the incumbent was relying on is the one AI erodes first. As I write, ElevenLabs and Anysphere’s Cursor are running on a small fraction of the headcount of the classical software leaders. Some of that edge is genuinely technical: ElevenLabs trains its own speech models, and Cursor’s harness is the product rather than a wrapper on someone else’s. But comparable models reach everyone eventually, and the providers will keep building better harnesses for all of us. What no incumbent can buy from a vendor is an organisation that was never taxed by a design built for a different world.

    What you can do from inside the box

    Most of the people reading this are the CTO, not the CEO. You can’t redesign the company. You can do something narrower and more useful. Pick one value stream where AI is meant to pay. Baseline the commercial number it is meant to move. Then map every approval and decision right sitting between an engineer’s idea and that number changing, and notice how much of the map lies outside your department. Take it to the person who owns the design, along with one 30-day experiment that needs exactly one decision from them. Either they make it, and you have found out the design can move; or they don’t, and you have learned more about your ceiling than another quarter of pilots would have told you.

    Where I could be wrong

    A good-faith explanation competes with mine. The technology is young, the data plumbing is bad, and most of the 2025 pilots were badly scoped experiments run by people learning on the job. On that reading the numbers improve on their own as the tools mature, and nobody has to touch the org chart.

    A second explanation is harder for me to dismiss. The causal arrow might run the other way: firms that found real value in AI then changed decision rights and reporting lines because the economics finally justified it, which would make an operating-model change the consequence of success rather than its precondition. The two readings differ on timing, so the test is whether the structural change came before the commercial result or after it. I haven’t found enough cases where both dates are public to settle it.

    If immaturity is the cause, results should climb with model capability while the design stays where it is. My prediction is that task-level speed keeps climbing and the business-level numbers stay flat, because the constraint has moved somewhere the tooling cannot reach. I would be wrong if you can show me an organisation that got a durable commercial result from AI while its decision rights, its incentive measures, its funding ownership and its cross-functional process ownership all stayed put, or a workflow redesign inside one department that moved a business number and held for 18 months. One such case and I have to narrow this. Several and it is wrong.

    Before publishing that offer I went looking myself, against four criteria fixed in advance: a realised commercial result on a P&L line rather than an adoption or task-time metric, held across four quarters, no reporting-line or incentive or new-function change in the same period, and no statement from the accountable leader describing a changed view. I found 18 public cases that came close enough to code. None met all four. That is weaker evidence than it sounds. Organisations publish what changed and almost never publish that they changed nothing and it worked, so the public record is far better at showing movement than at proving stillness.

    The wall you cannot delegate

    The change lives at the top of the organisation, in structure and rewards, and at the deepest level of engagement, in trust and identity. Both of those are the leader’s own ground. That ground belongs to the leader, not to the coach, the transformation office, or the Chief AI Officer, whose written charter, in the 50 I coded, almost always stopped short of the operating model around the AI work. The wall you can’t pay someone to cross for you is the one that runs through your own assumptions about how the organisation should be built, and whether the goals you have been optimising for still fit the situation you are actually in.

    This is also why sponsor continuity matters, but departure is a warning sign rather than a law of organisational physics. A design is exposed when its decision rights, rewards and governance still depend on the person who sponsored it. The practical test is whether the new design has become how the organisation governs itself before that person goes. Craig Larman has said, in the LeSS courses I have sat in, that he has never seen an Agile transformation outlast a change of leadership. I take that as a severe practitioner warning and no more, because I tried to test it and couldn’t: a first pass found reversion signals far more often after a sponsor left than when one stayed, but a properly matched replication failed its own coverage rules. So the status is a risk worth auditing, not a mechanism I can show you.

    Deming is supposed to have answered executives who asked him to train their people: come yourself, or send no one. Robert Kegan and Lisa Lahey give the mechanism in Immunity to Change. A paradigm becomes available to change only when it moves from something you are to something you have. Turning that on your own paradigm is the hard part, because the thing doing the looking is the thing you are trying to look at, which is why it so often takes someone outside the pattern to help you see it.

    This is the argument underneath a book I have been serialising, Come Prepared to Die. I won’t rehearse it here. The short version is that the structural case and the inner one are the same case seen from two sides. The person who owns the design can’t delegate the work of examining the assumptions behind it, and no title, and certainly no model, can do that work on their behalf. For the change to outlive them, its decision rights, rewards and governance also have to become properties of the organisation rather than permissions rented from one sponsor.

    The uncomfortable question

    If you are standing up an AI transformation right now, ask a narrower and more uncomfortable question than which tools to buy. What is your organisation actually optimised for? Who holds the authority to change that? And does the person holding it know what you know? If the effort is once again boxed into the delivery corner and handed to someone without the remit, then you already know how it ends.

    I work with engineering leaders on exactly this kind of paradigm work, the deeper the better. If it’s live for you, I’m happy to talk: schedule a 30-minute virtual coffee at hi.chrisgagne.com.

    Some book links here are Amazon affiliate links; if you buy through them I may earn a small commission, at no cost to you.

  • grounded-forge, by example

    grounded-forge, by example

    I’ve open-sourced grounded-forge, a working example of grounding an AI assistant in sources you trust. This post introduces it through one session, receipts included.

    grounded-forge is my answer to the problem in the last post: an AI assistant defaults to the most-published take, not the best one, because the consultant layer outwrites the original thinkers and the models inherit the ratio. Ask a stock model about mission command and you get the LinkedIn consensus on mission command. The fix is to take source selection back.

    The tool is open source. You pick the material you’d stake a real decision on. A model reads each source in full under a structured nine-pass protocol that checks every quotation verbatim against the text, then projects the source onto the tasks your assistant serves, once, at ingestion time. At runtime the assistant answers from that pre-audited material; the expensive, error-prone reshaping is already done.

    What follows is the transcript of one session, working notes first, then the essay. I want to be plain about authorship: the transcript in the card below was generated by the model, and it matches the copy that ships in the repo; apart from trimming the lens list to the four the public repo ships, I have not touched it. The slate notes labelled “Chris” are mine, added afterwards to narrate what the system is doing at each step. My contribution otherwise is the question, the choice of sources, and the system underneath, which is built to ground every load-bearing claim in a source passage and to mark verified quotations as it goes.

    The corpus behind the session is the demo library that ships with the repo (27 sources at the time), among them MCDP 1: Warfighting, the US Marine Corps’ capstone doctrine (public domain), and the OpenStax business textbooks, the standard management curriculum (openly licensed). My prompt was one line: “Please draw some interesting parallels between US Marines war-fighting doctrine and modern ‘textbook’ business administration.”

    The [V] markers in the essay mean verbatim: checked against the ingested source text by a source-only audit at ingestion, before any question existed. (The auditor is itself a model working under protocol; no human read every line. What the marker buys you is a specific file and a specific passage to check when a claim looks wrong.) The trace at the bottom names which projections were read in full, and the closing section declines the one transfer the sources themselves don’t support. The transcript itself, without my notes, ships in the repo as a captured session. Where a [BT] appears, the idea is borrowed through the source from an author the corpus does not hold, and labelled as such.

    Source-grounded synthesis · Protocol S (Synthesis) · 2026-07-23 · Claude Fable 5

    USER · /answer-from-corpus

    Please draw some interesting parallels between US Marines war-fighting doctrine and modern “textbook” business administration.

    ASSISTANT

    I’ll answer this by drawing on the corpus. First, let me classify the query shape and check for applicable lenses.

    Shape: This is a Synthesis question — it asks what the corpus says across two domains (Marine warfighting doctrine, textbook business administration) and wants parallels drawn between them. Breadth across sources is the point.

    Chris: It has detected this is a synthesis query. That is one of three shapes it routes: a named lookup reads one source and stops; a diagnostic maps your situation into one task domain; synthesis wants breadth across the corpus. The shape decides which indexes get read and in what order.

    Tool calls – Read lens-index.json (full) — four lenses now: builder, pm-bounded-by-ba-role, agentic-builder, chris-gagne-consultant-coach.

    Lens-applicability check: Lens-neutral. The query asks for an intellectual synthesis between two doctrines — not a role-bound artefact, not a reader-typed deliverable, and not an application of the named-person lens. No lens applies.

    Chris: It has decided not to apply a lens, as none fits the question. A lens reweights what a distillation surfaces for a particular role (a CTO and a builder need different salience from the same source), and the check runs before decomposition, while it can still shape the sub-claims. It declines the four on offer rather than forcing one.

    Sub-claims (the parallels to test):

    1. Nature of the decision environment — friction and uncertainty vs bounded rationality and scarcity
    2. The decision cycle — OODA vs the textbook multi-step process
    3. Who decides — mission command vs centralisation/decentralisation
    4. Priority-setting — Schwerpunkt vs the textbook prioritisation apparatus
    5. Error asymmetry — inaction vs commission
    6. Doctrine as thinking vs framework as tool
    7. Limits of transfer — the adversarial Zweikampf frame

    Tool calls – Read slug-table.json (full) — 27 sources; 00q = mcdp1-warfighting, plus the OpenStax family. – Read concept-index.json (full; exceeded the single-read cap — two paginated reads, 4,210 lines). – Read distillations/decision-making/task-index.json (full; two paginated reads, 1,809 lines).

    Chris: This is the routing pass, and the reads are whole-file by rule. The concept index is bigger than a single read allows, so it paginates to the end rather than grepping for a likely key. Grep answers the question you thought to ask; a full read is what lets the index volunteer sources you didn’t think to ask for, which is exactly what happens next.

    Routing findings. The concept index carries a dense MCDP-1 concept family (ooda-loop, mission-tactics-auftragstaktik, intent-two-levels-up, schwerpunkt-main-effort, leniency-overbold-severity-inaction, zweikampf-violent-struggle) and routes the business side beyond the management shelf: theory-of-constraints and relevant-costs/sunk-costs live in 00b (Principles of Accounting Vol 2), and sunk-cost, opportunity-cost, and marginal-analysis concepts live in 00e (Principles of Economics 3e). One gap to report honestly: the task index carries no rows for 00q at all — the situation router does not yet cover the corpus’s newest source. Routing for the Marine side therefore rests on the concept index and slug-table, which is sufficient here.

    Chris: The payoff of the full reads is in the findings above: accounting and economics surfaced for a Marines question, which no skim of filenames would suggest. And when the router itself had a hole (no task-index rows for the Warfighting source), the session reported the gap instead of papering over it. That is the no-silent-degradation rule doing its job; I fixed the index the same day.

    Candidate set and dedup. 00q (Warfighting), 00j (Principles of Management), 00h (Organizational Behavior), 00g (Introduction to Business), 00b (Accounting Vol 2), 00e (Economics 3e). 00c (Business Ethics) matched on stakeholder-theory for sub-claim 7, but 00h (Follett) and 00j (Freeman) already carry that sub-claim — skipped under the diversity-aware dedup rule.

    Tool calls – Read mcdp1-warfighting-decision-making.md (full). – Read openstax-principles-management-decision-making.md (full). – Read openstax-organizational-behavior-decision-making.md (full). – Read openstax-introduction-business-decision-making.md (full). – Read openstax-accounting-vol2-decision-making.md (full). – Read openstax-economics-3e-decision-making.md (full).

    Every sub-claim now has the Warfighting side and at least two textbook-side sources, with convergence and divergence both represented. The accounting and economics rows extend the comparison well beyond the management shelf. Let me write the essay.

    Chris: Six distillations read in full, not as retrieved chunks: the four any skim would find, plus the accounting and economics volumes the concept index surfaced. Each is a source already projected onto decision-making at ingestion, quotes checked and marked in place. The costly work ran once, before any question existed. It also declined a seventh read: Business Ethics matched, but the stopping rule found the sub-claim already carried.

    Two Doctrines of Deciding Under Uncertainty

    Read MCDP 1: Warfighting against the OpenStax business shelf — not just the management row, but the accounting and economics rows beside it — and the same question keeps surfacing from both sides: how do you decide well when you can’t see the whole board? The parallels are real and useful. So are the divergences. And one claim the narrower comparison invites turns out to be wrong once the whole shelf is on the table.

    Where they agree: the constraint is permanent

    MCDP 1 names the environment before naming any method: war is friction, uncertainty, fluidity, disorder — constitutive features, not exceptions. The doctrine is blunt that “no degree of technological development or scientific calculation will diminish the human dimension in war” [V]. Decision frameworks that assume the world is reducible will fail, because the world isn’t.

    The management texts arrive at the same place through cognitive science. Bounded rationality anchors both OpenStax management volumes: “for complex issues we cannot be completely rational because we cannot fully grasp all the possible alternatives, nor can we understand all the implications of every possible alternative” [V]. And the economics text arrives through a third door, defining its whole subject as “the study of how humans make decisions in the face of scarcity” [V] — adding that “the ultimate scarce resource is time — everyone, rich or poor, has just 24 expendable hours in the day” [V]. Three disciplines, one founding condition: the limit on the decider is permanent, so you decide inside it. Introduction to Business concedes the same about control: “managers are primarily adapters to, rather than agents of, change” [V]. Neither tradition believes the decider owns the field.

    The decision cycle: same skeleton, different clock

    MCDP 1 credits Boyd for Observe–Orient–Decide–Act and stakes its claim on tempo: the side that completes the cycle faster, or that the adversary cannot orient to, wins. The textbooks run the same loop at lower velocity — recognise, generate alternatives, analyse, select, implement, evaluate in the management volumes; a five-step variant closing with feedback in Introduction to Business, which warns that managers can get dangerously stuck at any stage. Analysis paralysis is a slow OODA loop by another name.

    The first-order emphasis gap is real: the management version optimises the quality of each step, the Marine version the speed of the whole loop relative to the opponent’s. But the economics row complicates the neat contrast, because it prices delay explicitly: “the opportunity costs of waiting time can be just as important as costs that involve direct spending” [V]. The textbook shelf does carry a clock; it lives in the economics text rather than the management one, and it prices your waiting against your own alternatives rather than against an adversary’s cycle.

    Who decides: mission command in three vocabularies

    MCDP 1’s answer to who should decide what is mission tactics: the senior gives the what and the why — the commander’s intent — never the how, and intent must be understood two levels up. Decentralisation works only when subordinates hold the intent; without it you get incoherence, not empowerment.

    Introduction to Business debates the same trade-off as centralisation vs decentralisation — warning both that centralisation can prevent quick local decisions in dynamic environments and that decentralisation without skills or training can produce costly mistakes — and renders it structurally as organic versus mechanistic design. The accounting volume adds a third vocabulary the narrower read missed: responsibility centres, which align decision authority with information access and accountability. That is the information-logic of mission command in accounting dress: push the decision to where the information lives, and hold the decider accountable for what they control. What the textbook shelf still lacks is MCDP 1’s sharpest instrument — intent two levels up as the specific content the empowered subordinate must hold. The textbooks say decentralisation needs skills; the Marines say precisely which skill: the boss’s boss’s purpose.

    Priorities: the textbooks have a Schwerpunkt after all

    Here the whole shelf corrects the essay a narrower read produces. Compare MCDP 1’s Schwerpunkt — name one main effort; everything else supports it; supporting yields when they conflict — with the management shelf alone, and the textbooks look like better-prioritisation people: Drucker’s eight goal areas, SWOT, weighted analysis. Prioritising better preserves the multi-priority frame; naming a main effort breaks it.

    But managerial accounting carries the textbook tradition’s own one-filter discipline: constrained-resource allocation. When a resource binds, rank every product by contribution margin per unit of the constraining resource — not by unit margin. The highest-margin product is often the wrong priority; what matters is yield against the single binding constraint. The volume carries this from Goldratt’s Theory of Constraints [BT], and it does to product-mix decisions exactly what Schwerpunkt does to operations: one thing governs, everything else is ranked against it.

    The residual difference is sharper for being narrower. The Marine main effort is chosen — a bid, an act of will that can be shifted mid-fight as gaps appear. The accounting constraint is discovered — an empirical fact about capacity you find and exploit. One tradition names its focus; the other finds it. That is the true divergence, and it is smaller than the one the management shelf alone suggests.

    The asymmetry on inaction: the mirror has a crack

    MCDP 1 rejects zero-defects outright: severity on errors of inaction, leniency on overbold errors. The management texts brake in the opposite direction, and for cause: their catalogued failures are failures of commission — escalation of commitment, “the tendency of decision makers to remain committed to poor decision, even when doing so leads to increasingly negative outcomes” [V], groupthink [V], and the remedies are devil’s advocates and precommitted monitoring.

    Yet the economics row again refuses the clean mirror, because it carries anti-hesitation machinery of its own. It names why waiting feels safe: loss aversion — a loss pains roughly 2.25 times more than an equivalent gain helps [BT] — and it orders the release: “the lesson of sunk costs is to forget about the money and time that is irretrievably gone and instead to focus on the marginal costs and benefits of current and future options” [V]. The accounting volume drums the same refrain into every short-term decision category. So the honest contrast is institutional, not intellectual: the textbook shelf hands the manager tools against both failure directions and leaves the calibration to judgment; the Marines legislate the asymmetry as policy, because their characteristic death is hesitation. The tools exist on both sides. Only one side turned them into doctrine.

    Doctrine as thinking: three-way agreement

    MCDP 1’s most general claim — doctrine is a way of thinking, not a checklist — turns out to be the one the whole shelf endorses. The management texts warn against running frameworks mechanically and insist a good process can yield a bad outcome under uncertainty. The accounting volume makes the same move in its own material: costs are not intrinsically fixed or variable — the classification is decision-relative, and the practitioner’s question is “fixed for what decision?” A label that must be re-derived per decision is judgment wearing a taxonomy. The framework is scaffolding; judgment is the thing. On this, no daylight.

    The one parallel that doesn’t hold

    MCDP 1 frames war as Zweikampf — a violent struggle between two hostile, irreconcilable wills — and its edge follows from that frame: tempo, surfaces-and-gaps, getting inside the opponent’s cycle. The Warfighting distillation itself flags this as the piece that usually does not transfer. The textbook shelf encodes the opposite cosmology three times over. The management row offers Follett’s integration — both parties reaching the creative agreement that satisfies both [BT] — as the strongest conflict outcome, and flatly rejects harming stakeholders for shareholder return: “increasing the wealth of shareholders is not an acceptable reason for causing harm to others” [V]. And the economics row is built on the founding scene of mutually beneficial exchange: comparative advantage says the counterparty is a trading partner whose different opportunity costs are the very reason to deal, not a will to break.

    So the synthesis holds, sharpened: the machinery transfers — OODA, mission command, constraint-governed priority, the case against hesitation, doctrine-as-judgment — and much of it has closer textbook cousins than a narrow read reveals. The cosmology does not transfer. Reach for the Marine playbook when there is genuinely an adversary and a clock; reach for the shelf when the room holds stakeholders and trading partners, which for most of management is most of the time. The mistake that survives every widening of the corpus is the same one: using the Zweikampf frame on people who were never your enemy.


    Trace [Synthesis, Depth-over-breadth, lens: none, corpus: demo]: Pass 1 (lens-index + slug-table + concept-index (2 paginated reads, 4,210 lines) + decision-making/task-index (2 paginated reads, 1,809 lines); gap noted: task-index carries no rows for 00q) → sub-claims: environment (00q + 00j/00h bounded-rationality + 00g + 00e scarcity); cycle (00q + 00j/00h six-step + 00g five-step + 00e waiting-cost); decentralisation↔mission-command (00q + 00g + 00b responsibility-centres); priority/Schwerpunkt↔constrained-resource (00q + 00j + 00b ToC-[BT]); inaction-asymmetry (00q + 00j/00h escalation + 00e loss-aversion/sunk-cost + 00b sunk-cost); doctrine-as-thinking (00q + 00j + 00b decision-relative-costs); Zweikampf-limit (00q + 00j Follett/shareholder-rejection + 00e comparative-advantage) → dedup: 00c matched (stakeholder-theory) but skipped, 00h+00j already carry the sub-claim → Pass 2 (6 distillations read in full: 00q, 00j, 00h, 00g, 00b, 00e) → verbatim [V] markers surfaced in-band from the distillations. Lens-neutral: no role-bound reader, no lens applied (4 lenses checked).
    Chris: The trace is the receipt: protocol, corpus, the lens ruling, which sources fed each sub-claim, what was read in full, and even the read it declined. It is what lets you audit the answer instead of taking its fluency on trust. A trace that names what was read shows what wasn’t.

    That’s the whole exchange, trace and all. The build system, the 9-pass ingestion protocol, and the evaluation showing where the architecture wins, and where a naive corpus read beats it, are open at github.com/chrisgagne/grounded-forge.

    If your work leans on sources you actually trust (doctrine, standards, your own case notes), the pattern transfers: each source is read once under audit and projected onto the tasks you repeat, and answer time becomes a lookup. I run the same machinery for After-Action Review and retrospective facilitation. My hunch is that the audit trail matters more than the essay. Fluent prose from a model is cheap now; what I wanted was prose I can check.