How software organisations actually deliver: team-of-teams structure, flow economics, cost of delay, and the constraints that decide throughput. Reading the system, not blaming the team.
You defended a decision after it failed. The decision was a hypothesis, and a hypothesis that fails is data. When a decision fails, reading the result and protecting yourself become two jobs competing for the same moment, and it is easy to mistake the second for the first.
Chapter 9 made outcomes the unit of management. This chapter asks what happens when the outcome misses. Does the leader defend the decision or ask what it taught them?
Pole A: attached to outcome. The decision must yield this result. If it doesn’t, the decision was wrong.
Pole B: attached to inquiry. The decision was a test of a hypothesis. The outcome is data. The question is what to learn from it.
This doesn’t reverse chapter 9. Outcomes are still the unit of management; a missed outcome is exactly the signal that the decision was a hypothesis. Pole A’s error isn’t caring about the outcome. It is fusing the decision’s rightness—and the leader’s—to whether this particular outcome landed.
When a decision fails, it is easy to mistake the second job for the first. The if-only counterfactual after Amy Edmondson, Right Kind of Wrong; the fork is the chapter’s own.
Where Pole A is right
Pole A is right when accountability requires defending the commitment. That may mean contractual delivery against a signed scope, a public commitment to investors with material implications for share price, or a regulatory commitment where non-delivery could cost the organisation its operating licence. Some crisis communication demands the same discipline: the leader may need to hold a position in public for months while the team works on the underlying problem.
The leaders I have worked with in those domains—myself among them—had good reason to default to Pole A. Defending the commitment was their job, and the discipline they built doing it made them trustworthy enough to lead where keeping a commitment mattered. That same discipline is what built the relationships with boards, investors, regulators, and key customers where a senior leader now has standing, and none of it was waste.
Pole B depends on that trust. The senior leader doesn’t need to apologise for having earned it.
Where Pole B is right
Pole B is right when the decision was a hypothesis dressed up as a commitment, or when defending it is doing more for the leader’s identity than for the outcome. The second case is harder to see from inside because defending commitments under uncertainty is a discipline most leaders have built over a career.
Most decisions made in complex conditions under uncertainty are hypotheses. The leader makes them with limited information and under time pressure, while the team still needs something definite to work towards. Accountability often demands the language of commitment because that language helps the team coordinate. The leader who speaks commitment in public can still be running inquiry in private.
The leader who speaks commitment in public can still be running inquiry in private.
Once the leader treats the decision as settled, inquiry becomes much harder.
Counterfactual thinking as identity-protection
Amy Edmondson, in Right Kind of Wrong, names the cognitive habit underneath the Pole A response, drawing on a classic Olympic-medallist study. Bronze medallists framed their result as a success (a medal at the Olympics!), while silver medallists framed theirs as a failure relative to gold through what psychologists call counterfactual thinking.
The counterfactual governs whether an outcome is even felt as a failure. The silver medallist’s if only is Pole A’s emotional signature.
The silver medallist’s if only is Pole A’s emotional signature. The Olympic-medallist study, after Amy Edmondson, Right Kind of Wrong.
Edmondson calls the self-flattering version self-serving analysis: the mind generates alternative pasts where the leader was right, and generating those pasts protects identity rather than learning from the result.
Counterfactual rumination has a recognisable form: if only X had gone differently, the decision would have worked. I told myself a version of it for years. In 2014 I published a list of recommended Agile reading, put Jeff Sutherland, Scott Downey and Björn Granvik’s shock-therapy paper on it, and noted that I’d prescribe the same thing: adopt Scrum 100% from the start. Going back to it now, the paper is stricter than I had remembered. The coach sets the rules, and a team earns the right to change one only after “three consecutive, successful Sprints”, a “240% increase in Velocity” and a business reason every member agrees to; the Scrum framework itself never changes. When a team fell short, the paper found the cause in enforcement: the one low-performing team was the one whose Scrum Master “failed to enforce constraints”. The belief that produces is closed. A team that broke through had proved the rules; a team that hadn’t hadn’t yet earned the right to question them. Nothing a team did could count against it, and that took me years to see. A leader practising inquiry can stay with a harder account: the decision didn’t work; here is what it taught us; here is what we will hold differently.
Both responses look similar from the outside. They diverge in what happens next: one goes looking for evidence, the other for a version of events in which the call was right.
Edmondson’s four criteria
Pole B requires more than openness or humility. Edmondson teaches how to fail well. An intelligent failure meets four criteria:
It enters new territory: no one had the answer in advance.
It serves a goal worth the attempt.
It is informed by available knowledge: you had reason to believe the hypothesis could be right, and the failure isn’t one your homework would have caught.
It keeps the risk as small as possible while still producing meaningful learning.
A bonus: the learning is mined and shared, so the failure doesn’t recur.
Edmondson separates intelligent failure from error: “Error implies there was a ‘right’ way to do it in the first place. Intelligent failures are not errors.” The Pole A leader who treats every failed decision as an error to be explained is using a vocabulary that doesn’t fit the work.
Edmondson’s failure taxonomy has three types:
Basic failures are caused by error in known territory.
Complex failures occur when multiple causes line up, warning signs get missed, and a recovery window closes unused.
Intelligent failures are the ones that clear those four criteria.
Every failure is worth learning from; inquiry isn’t reserved for the intelligent kind. What the criteria decide is narrower: whether the failure was worth running, or whether it is an avoidable mess dressed as a noble experiment.
The four criteria (plus the bonus: the learning is mined and shared) after Amy Edmondson, Right Kind of Wrong; the three examples are the book’s own.
Toyota Kata: problems are jewels
Mike Rother’s Toyota Kata puts Pole B into daily lean-improvement practice. The improvement kata runs through PDCA cycles in which each obstacle becomes the next thing to learn from. “Problems are jewels” is the Toyota colloquialism: the problem is the curriculum.
The mentor asks the mentee five questions:
The five coaching questions after Mike Rother, Toyota Kata.
What is your target condition?
Where do things actually stand right now?
Which obstacles are now in the way of reaching the target condition, and which one are you tackling next?
What will your next step be?
When can we go and see what that step has taught us?
Question 5 is the inquiry move. The next step is an experiment with a specified learning horizon, and the next coaching cycle starts by examining what was learned.
The kata turns Pole B into a daily practice the team holds together. Inquiry no longer depends on one leader’s capacity to sustain it alone.
Embrace change, then pivot
Pole B has a precondition, and it is bigger than it first looks. The engineering half is familiar: a team that can refactor, pair, and integrate continuously doesn’t need to defend yesterday’s design when today’s evidence contradicts it. Continuous integration makes the technical pivot cheap. AI-assisted refactoring and code generation push the cost of change lower still.
But cheap-to-change code isn’t enough to pivot a strategy, and the harder precondition is organisational. To pivot is to abandon work already in flight. A team running several epics at once to keep everyone busy—the high-utilisation trap from chapters 4 and 7—can’t pivot without writing all of them off at the same time, so it defends the plan instead. The team carrying one small bet can drop it and turn. Low epic-level WIP is what makes a pivot affordable; high WIP forces Pole A even when the leader wants Pole B. The context keeps the team shipping yesterday’s answer after today’s evidence has disproved it because the sunk cost is too high to accept.
A pivot follows validated learning: we tested a hypothesis, the evidence came in, and the next hypothesis follows from what we learned. A board is more likely to absorb that when the evidence, the learning, and the next test are all on the page. An unexplained reversal is a harder sell.” The difference isn’t that the words are prettier. A pivot names a decision still under inquiry; abandonment names a retreat.
A pivot names a decision still under inquiry; abandonment names a retreat.
This is the chapter 9 move at the strategic level: we named the bet, said what it should produce, read the result, and named the next bet. The board still gets its number, the same trade chapters 7 and 9 made.
The protector part’s job
When a failed decision threatens a leader’s identity as competent, what Internal Family Systems (Schwartz) calls a protector part—here, the Defender—generates counterfactuals and attributes the result to factors outside the leader’s control. The leader appears confident. Inquiry has stopped.
The Pole B move is to recognise the Defender part, name what it is defending, and ask it to step back and give space so the rest of the leader’s judgement can engage the evidence. That is the work the next chapter—Just This—takes up directly. A leader who is actually confident, sure of the value they bring, can be wrong now and then and interrogate their own decisions without waking the Defender.
A representative case
Picture a Series A founder who has committed publicly to a particular customer segment (the commitment is in the deck the investors funded). Six months in, customer-acquisition cost in that segment is running well above the model’s projection and the success metrics are trending the wrong way, while the team has begun landing customers in an adjacent segment where the numbers are far better.
Across a series of leadership meetings, the founder defends the original selection with increasingly sophisticated arguments. By the third meeting the arguments are more sophisticated than the numbers are good. That reads two ways around the table: rigour about a hypothesis the team hasn’t properly tested, or a leader defending a choice. The founder holds the first reading.
In this case the intervention comes from outside. In a one-on-one, an outside coach asks, what would it cost you to publicly update the strategy? The real answer names the cost: the founder’s identity as the leader who picked the right thing, and the investors who will read the update as the founder was wrong. Over several weeks, the founder works through that cost and updates the strategy in a board memo, describing the change as a pivot informed by six months of validated learning. The pivot language gives the board a way to absorb the change.
The board’s response surprises the founder. They are relieved; they had been wondering when the founder would name what the metrics had shown for two quarters. In Edmondson’s terms, the founder nearly squandered a recovery window.
Honour what was genuine in the original defence. The founder was also doing part of the job well: holding the strategy through enough friction to test it. A genuine commitment and the Defender were operating on top of each other. Updating both lets the founder enter the next decision with judgement and identity intact.
The diagnostic move
Three questions for last Tuesday’s post-mortem.
Which pole was I claiming? In the language I used about the failed decision, did I describe it as a hypothesis we tested, or as a commitment we kept (or failed to keep)?
Which pole did my response show? When the evidence came in, did I generate alternative pasts where the decision would have worked, or did I stay with what the evidence taught us?
Which pole does the work require? If the accountability frame genuinely requires defending the commitment in public, Pole A is defensible. If the public commitment is being used to keep you from treating the decision as a live hypothesis, the Defender is running.
The failed-decision post-mortem is where this choice bites hardest at the top. If you sit in the CTO or VPE chair, have the private conversation with the CEO. Bring one specific decision from the past quarter that didn’t work and ask, not what went wrong but what hypothesis were we testing, and what did we learn?
Ask it in private first. If the CEO can work through the learning in a one-on-one, they can carry it into the board conversation.
If the Defender is running strongly, drop the diagnosis and ask the question that gives it space: what would make updating this strategy safe to say out loud? Sometimes the answer is data; sometimes it is a sponsor’s backing, a staged disclosure, or wording that keeps the legitimate commitments intact. The sentence your CEO can carry to the board is, “We learn at the rate we can say a decision didn’t work.”
We learn at the rate we can say a decision didn’t work.
Chapter 7’s optimum failure rate is the aggregate version of this: cheap experiments should fail about half the time. A team whose bets almost never fail isn’t excellent; its experiments are confirmations, and confirmations teach little. A failure rate near zero is Pole A in the language of excellence.
The exercise
Apply the Toyota Kata five questions to your next stuck decision. Pick a current open question where you have been defending the position you committed to two months ago, then walk through the questions with a coach or peer who isn’t invested in the answer: target condition, actual condition, obstacles, next step, and the question that does the work, when can we go and see what we have learned?
If you can’t specify a date by which the next step’s learning will be visible, you are still running Pole A. The Pole B answer specifies the learning horizon.
A companion move is the decision audit. Take a recent decision that didn’t work. Did you defend the decision or refine the model? Name your response out loud to someone you trust enough to disagree with you. The harder, identity-level work starts there.
Going upstream
In-text: Amy Edmondson, Right Kind of Wrong, on the four criteria for intelligent failure, the three failure types, and the counterfactual-thinking pattern; Mike Rother, Toyota Kata, on problems are jewels and the five coaching questions.
Also touched: The protector named here as the Defender is Internal Family Systems language (Richard Schwartz, No Bad Parts); the Just This instalment works it directly.
Go deeper: Kent Beck, Extreme Programming Explained, on embrace change and the continuous-integration discipline that makes pivoting cheap; Eric Ries, The Lean Startup, on the pivot and validated learning; the overproduction the Lean tradition names (Taiichi Ohno through Beck and Ries) for inquiry shipped on top of an architecture that can’t support it.
Watch. Astro Teller, The unexpected benefit of celebrating failure (TED, 16 min). Kill-and-pivot stories from Alphabet’s X: bonuses paid to everyone on teams that chose to kill their own projects, from pairs to groups of 30+.
Ten chapters in, the tools have become something to grip, and the grip tightens with every one you add. Every framework in this book becomes a trap the moment you obey it instead of using it.
Chapter 11, Just This, asks you to notice that reflex before choosing a tool. The one you reach for first may already be telling you what to see, before you have met the situation actually in front of you.
I work with engineering leaders on exactly this kind of paradigm work, the deeper the better. If it’s live for you, I’m happy to talk: schedule a 30-minute virtual coffee at hi.chrisgagne.com.
Some book links here are Amazon affiliate links; if you buy through them I may earn a small commission, at no cost to you.
I first published a version of this piece on LinkedIn in 2021, when the argument ran on Weinberg’s table and Kniberg’s prioritisation illustrations alone. They carry the “don’t do everything at once” point well enough. What I have built since is a working cost-of-delay tool, grounded in Don Reinertsen’s Principles of Product Development Flow—the source of both the cost-of-delay economics and the WSJF / CD3 prioritisation heuristic—and the practical cost-of-delay work of Joshua Arnold, who did as much as anyone to turn it into something teams actually use. The “cost of delay” is no longer a rhetorical lever for me: the tool ranks every story by it, and where the business states what the work is worth, it can put an estimated dollar figure on the delay. Also, the Cost of Delay work has shown me that most of the value comes from focus, not necessarily getting the prioritisation perfect.
Gerald Weinberg’s book “Quality Software Management: Systems Thinking” is more than 30 years old. While it’s not one of the most highly-read and recommended classics in the adaptive canon, it’s the source of a frequently quoted table of data:
Visualised as a graph, the waste caused by context switching really stands out:
Weinberg’s figures, Table 2-1. Each coloured block is one project’s share; grey is time lost to switching.
Worse, the waste caused by project switching isn’t due only to the losses related to cognitive overhead. Failure to prioritise often leads to less revenue and even building the wrong product.
(Credit where credit is due: I first saw a version of the following illustrations presented by Jeff Sutherland, who adapted it from Henrik Kniberg. I’ve created new illustrations and expanded on them a bit. Kniberg’s own video on this exact topic is very much worth a watch.)
Let’s suppose your company wishes to ship three products: A, B, and C. To ship a product, a development team must complete tasks 1, 2, and 3 corresponding to that product:
Most companies are not very good at prioritising work. As a consequence, the prevailing belief is: “Everything is important. Get started on everything immediately!” The traditional delivery timeline might look like this:
A blue tag marks the day a product ships.
Depending on your software and how you break down your work, your roadmap probably looks like:
A blue tag marks the day a product ships.
Either way, notice that we are interleaving each task and that products A, B, and C are ready at roughly the same time, several months after we started.
The adaptive approach is very different. Instead, we proactively prioritise the work and focus on limiting our work in process:
A blue tag marks the day a product ships.
There are at least three significant advantages to this adaptive approach: lower cost, more revenue, and better product/market fit.
Lower cost of development
We lose 20% or more of our productivity in the traditional approach due to context switching waste. In this example, the company could go about twice as quickly if they switched to developing one product at a time instead of three.
Grey: time lost to switching between products.
More revenue, earlier
The switching waste is the smaller loss. Because we have nearly finished products B and C before finishing A, we had to do nearly three times the work (not including the context-switching waste) before we could ship Product A (in late May). Had we prioritised and focused, we would have been able to ship Product A in early February. Had we done so, we may have been able to collect revenue and feedback from our customers starting nearly four months earlier. In fact, the savings from not context-switching between projects may mean that we could ship products A, B, and C before we would have been able to ship just A in the traditional model.
Hatching: a product waiting between its tasks. Blue: a shipped product earning.
Better product/market fit
Products are rarely independent, and building the wrong product can be costly. In this example, we believed before starting the work that the customer needed Product C. We took the adaptive approach and built A (more valuable and/or cheaper) first. When we delivered A to our customers in early February, we learned that they liked it a lot and didn’t need C after all. Instead, they wanted us to work on D. We took February to finish most of B and get the initial prototypes for D built.
Dashed: planned, never built. Blue: a shipped product earning.
We can see that prioritisation and focus serve everyone: developers, stakeholders, and customers.
Good prioritisation is not a zero-sum game
Further, good prioritisation is usually not a zero-sum game. Many stakeholders will argue aggressively for their product to be worked on right away (and it’s no surprise: they may even have a bonus tied to it being completed by a certain date). However, the waste caused by project switching and the cost of delay is so significant that failure to prioritise may make all stakeholders worse off. In this case, the stakeholders behind projects A, B, and C agreed to prioritise based on the cost of delay and all were better off than if they had insisted that everyone’s work be tackled at once.
The same three epics, priced
Five years on, this argument became software. The infographic below, from my Delivery Intelligence work, prices the same scenario with cost of delay: keeping all three epics in flight collects $42.50 of value by the end of July; finishing one at a time collects $138.75 in the worst order and $176.25 in the best. In this toy scenario that is roughly 3–4× the value, and the ordering mattered far less than the focus.
Each epic needs three large tasks. Once shipped, A earns $10 a month, B $5 and C $20. Interleaved, each task runs about 60% longer: at three projects in flight, roughly 40% of the time is lost to switching. Relatively speaking, the order made little difference; the value came from the focus. Interleaving against one-piece flow, after Henrik Kniberg.
Feel it in five minutes: the character factory game
You don’t have to take my word for any of this, or Weinberg’s. Grab a pen, a sheet of paper with three columns, and a one-minute timer. Three customers each want a complete set of 25 characters—letters, numbers, vowels—and nothing partial counts. Round 1: keep all three orders moving by writing one character for each customer in turn. Round 2: finish one customer’s set before starting the next. One minute each. (The game is Henrik Kniberg’s multitasking name game, with a modification I learned from Monica Yap and a further modification of my own.)
The card assumes the same writing speed in both rounds. Focus still wins—two complete sets against none—because a complete set is what a customer can actually use. Run it live and a second effect appears: your total character count drops in round 1, because every switch costs a beat. That drop is Weinberg’s table on one sheet of paper, and the delivered sets are the revenue argument in miniature. Try it with your team this week.
Some book links here are Amazon affiliate links; if you buy through them I may earn a small commission, at no cost to you.
Russell Ackoff liked to “build the best car in the world” as a teaching example. Take one of every make sold in America, he said, and have your engineers pick the best of each part: the best engine, Rolls Royce; the best transmission, Volvo; the best alternator, Porsche; the best of every component a car needs. Now bolt those best parts together into one automobile. You wouldn’t get a functional automobile.
The performance of a system isn’t the sum of the performance of its parts. Russell Ackoff’s teaching example.
The performance of a system isn’t the sum of the performance of its parts. It is the product of how those parts interact, which means you can improve every part on its own and wreck the whole doing it with far less effort or neglect than you might have expected.
Eli Goldratt put it in a sentence: a system of local optimums isn’t an optimum system at all; it is a very inefficient system. The way most senior leaders structure improvement work—every department on its own metric, optimised in parallel, the whole supposedly improving as the sum—is the wrong move in any system whose parts share constraints.
Plan as hypothesis was chapter 7’s move. This chapter asks where in the system that hypothesis should concentrate, because when functions share a constraint, parallel optimisation can reduce total throughput.
🎧 Prefer to listen? This chapter is narrated in my own voice with ElevenLabs on Spotify (~28 minutes). Listen on Spotify.
The two poles
Pole A: local optima. Improve each department on its own metrics. The whole will improve because it is simply the sum of the parts. Run the function-level reports; reward the function leaders who move their numbers; trust composition.
Pole B: system constraint. Identify the one constraint that limits the whole. Exploit it. Subordinate everything else to it. Decide what will raise throughput, and protect it. There is a greater wholeness that is more than just the parts; it emerges when those parts interact.
Pole A is right in some contexts, while Pole B is the operational discipline most engineering leaders need to recover. Both depend on one deeper claim: a constraint helps organise the system.
Where Pole A is right
Pole A is right only when the system genuinely decomposes. That includes independent product lines with no shared engineering and separate geographies running their own supply chains. Decoupled value streams with no shared bottleneck resource also qualify. The clearest case is low-variability work run in independent operations whose bottlenecks don’t interact, so optimising each step in isolation adds up to a better whole.
Most senior leaders inherit Pole A from a time when their mandate was smaller and easier to decompose, and the local metrics genuinely did roll up. The reports were true then. The discipline you built while running departmental P&L reports got you here; none of it was wasted. The question is whether that discipline still fits the system you run now.
In contemporary knowledge work (most software, product development, and cross-functional services), operations rarely decompose so cleanly. The output of resources and people passes through shared bottlenecks, as do queue capacity and management attention (even when the org chart hides them). Most systems that leaders expect to break into independent parts won’t.
Where Pole B is right
Pole B is right where events depend on prior events and demand varies. In Goldratt’s manufacturing context, improving local efficiency on non-bottleneck resources reduced total plant throughput by pushing more work onto the bottleneck and consuming time that couldn’t be recovered.
Deming’s funnel experiment reaches the same result for stable production processes. Tampering with common-cause variation increases the variance. An operator who adjusts the funnel after each ball lands produces a wider distribution than one who leaves it alone. Local correction makes the whole system less stable.
Pole A leaders set utilisation targets per function and ask whether each function is fully booked. Pole B leaders ask which one resource is governing how fast the whole system can move today, then protect it. Their rigour lies in subordination: they choose what can wait so the constraint can work.
CTOs and VPEs: a constraint that sits inside engineering is yours to work. One that crosses functions, or needs capital or a change to the org structure, needs CEO sponsorship. For that second kind, put the actual constraint in front of your CEO (the one resource, decision, or capability governing the speed of the whole engineering system), alongside the current investment split across functions. The CEO then authorises the subordination of non-constraints, usually by defending the Pole B investment against pressure from other function heads whose metrics suffer when their resources serve the constraint instead of their own dashboards. Your CEO can carry that choice to the board in one sentence: “One constraint sets our speed; every other optimisation is spend that does not raise throughput.”
Goldratt’s five focusing steps
The Theory of Constraints has an explicit working procedure. First, identify the constraint. Second, decide how to exploit it, squeezing every useful hour from the bottleneck before adding capacity. Third, subordinate everything else to that decision, so non-constraints serve the constraint’s pace instead of their own. Fourth, elevate the constraint by adding capacity once exploitation is maxed. Fifth, return to step one because the constraint will move as soon as you fix the current one.
Goldratt wrote the warning in step five in capitals, and the retellings I read dropped it: go back to step one, but don’t allow inertia to become the system’s constraint. Yesterday’s exploitation rules and subordination policies can outlive the constraint they were built for. They harden into the new constraint.
The five steps force a CEO-level choice about which function’s metric is allowed to wait. In software development, I suspect the single-backlog move later in this chapter is the most consequential application. Step five is where inertia bites: the architecture review board that once caught real problems, kept after the teams learned to catch them earlier, becomes the queue every change now waits in.
The doctrine is right as far as it goes. It usually leaves out the people whose continuous adjustments make a constraint productive.
Juarrero: constraints are causes
Alicia Juarrero’s Dynamics in Action gives Goldratt’s claim its causal basis and changes what constraint means. Juarrero’s distinctive contribution to the philosophy of action is that constraints structure action and make it possible. In a hierarchical organisation, constraints link levels by closing off alternatives at one level and opening new ones at another.
The kind operations leaders need is Juarrero’s third: parts interact until the whole they form starts limiting and steering those same parts. A shared review queue is the everyday case. No one designs it as a governor, but once every team routes work through it, the queue’s depth begins deciding what each team can start, and the teams that created it now take their pace from it.
Goldratt’s constraint, the slowest step in the system, becomes the organising principle around which the rest of the system works. It changes what is possible. Subordination recognises that the constraint helps hold the system in coherent operation.
The plant’s bottleneck sets the rhythm for everything upstream and downstream. A hospital’s emergency department is where the rest of the hospital’s flow rates are settled. On an engineering team, a senior architect may be where open design questions are forced into closure. A leader who treats the constraint only as a restriction will optimise around it; a leader who treats it as a cause will build the rest of the system to support the work it does.
Treat the constraint as a restriction and you optimise around it; treat it as a cause and you build the system to support it.
People continuously create the structure
Juarrero’s account changes how I read the subordination step. The constraint matters because people continuously adjust around it. Practitioners hold its productive structure in place through moment-by-moment compensation for the conditions it creates. Stop that compensation and the constraint becomes a bottleneck with none of the work around it that made it useful.
An engineering manager who treats the constraint as a mechanical feature to optimise around will miss the operators’ continuous work of creating structure. The team may read that optimisation as a vote of no confidence in their judgement and scale back the small adjustments they had been making. The system then degrades in ways the manager can’t read until something visible breaks. The constraint didn’t get worse. The work around it did.
The constraint didn’t get worse. The work around it did.
Pole B asks leaders to identify the constraint and name what practitioners already do to make it productive, then protect that work from optimisation programmes that would interfere with it.
Snook: the fallacy of social redundancy
Scott Snook’s Friendly Fire shows how social redundancy can make a system unsafe. He studies the 1994 Black Hawk shootdown over northern Iraq, where U.S. F-15 fighters shot down two U.S. Army helicopters while an AWACS aircraft watched, killing 26. The AWACS had 19 crew members. “More people looking over shoulders than there were shoulders to be looked over.”
Snook calls the operating assumption the fallacy of social redundancy. Technical systems add redundant components to cut the chance of an accident, but the group dynamics of the shootdown led Snook to conclude that two controllers needn’t beat one, and four leaders can be worse than two. Technical redundancy assumes that components are independent. People are interdependent, so adding more of them can diffuse responsibility, Latané and Darley’s bystander mechanism named in a different domain 40 years earlier.
Pole A usually answers a near-miss with more oversight: reviewers, approvers, and escalation paths. Snook argues that this can make the system less reliable. Clear ownership exploits the constraint, especially when leaders protect the people who hold it.
Team-level backlogs are local optimums that cap revenue
A bottleneck step, a senior architect, and an AWACS crew all make the constraint visible as a resource, person, or team. I met another form every day for years without recognising it, until a short film named it for me: several backlogs where the product needs one.
The film is by Michael James, a Scrum trainer whose work I keep sending to clients. James traces the idea to Bas Vodde and Craig Larman’s work on Large-Scale Scrum (LeSS). When an organisation scales from one team to many, it usually gives each team someone to act as product owner, then asks that person to manage a team backlog and grow team output. James calls this deviation a “team output owner” because output is what the organisation rewards with its design, if not its platitudes.
Each team output owner orders a private list to deliver the most value they can see, and each is doing thoughtful, careful work. The system sets the trap. “Keeping separate team backlogs, separate lists obscures this problem,” James says. “Our team’s top priority is less important than the work other teams don’t even have time to start.” Vodde and Larman’s remedy in LeSS is a single product backlog, one Product Owner with real authority, and teams that self-select from the shared ordering.
Blue: shipped this sprint. Pooling uses team three’s slots for team one’s stranded $70, $60, and $50: the extra $150 was sitting below a line drawn by the org chart. The chapter’s worked example. The single product backlog after Craig Larman and Bas Vodde, Large-Scale Scrum; the team-output-owner pattern after Michael James.
Imagine three teams, each able to ship three items in a sprint, with each team working from its own backlog. Team one has items worth $100, $90, and $80 at the top, followed by $70, $60, and $50. Team two has $30, $20, and $10. Team three has three $10 items. With three owners and three lists, each team ships its own top three: team one delivers $270, team two $60, and team three $30. The organisation delivers $360. Every team is busy, and every team is shipping the best work on its list.
Pool those lists and choose nine items for the whole product. The top nine are $100, $90, $80, $70, $60, $50, $30, $20, and $10. The same nine slots of capacity now deliver $510, with the same people working at the same pace. Pooling drops the three $10 items that team three would have shipped and uses those slots for team one’s stranded $70, $60, and $50. The extra $150 was sitting below a line drawn by the org chart while team three shipped small change because small change was the best work it could see. Goldratt’s sentence describes the prioritisation layer as exactly as it describes the plant floor: a system of local optimums isn’t an optimum system.
I didn’t want to trust a toy example on its own, so I built the measurement into the delivery-intelligence tool and ran it against the live backlogs of three value streams, 25+ teams in all. For every story, it forecasts three completion-date ranges from the same points-over-velocity arithmetic, using three widths of pooling, each range running from a P70 to a P95 date rather than a single point. A Team ETA uses one team’s throughput. A Capability ETA uses every team in the capability, while a Value Stream ETA uses the whole value stream. The wider pool ships high-value work sooner for the same reason the pooled backlog beat the three private ones.
The tool also flags the failure directly. When a story’s Team ETA is meaningfully sooner than its Capability ETA, the tool marks it as a local optimum: fast on its own list, but sitting where the most valuable work isn’t. The mark fires more often than I would like. Each one is the same trade as the $360 that should have been $510, happening one story at a time. Where work is genuinely team-local, the Capability and Team ETAs agree and no mark fires, which is what makes the flag worth reading when it does.
Synthetic data, not a client screenshot. Each ETA is a P70–P95 range, tinted by that cell’s own forecast reliability (red where the range is too wide to trust). The ⚠️ on a story’s Team ETA is the local-optimisation flag: the Team P70 lands earlier than the Capability P70, so the story ships sooner on its own team while the shared Capability backlog would deliver more total value first.
The obvious objection to one list is that teams can’t pull each other’s work: different skills, different systems, different context. Larman and Vodde’s LeSS treats that as a condition to grow out of rather than a fact to design around, deliberately encouraging multi-learning and a higher proportion of feature teams, and treating single-specialist teams as a pattern to leave behind. What’s new since they wrote is that AI cuts the cost of the growing-out: an engineer picking up an unfamiliar service now has something that will read the codebase alongside them, explain the local conventions, and draft the first change. A team that can pull any high-value item is therefore the condition worth building toward. Teams don’t have to become identical tomorrow. They keep their skills, context, and sense of a shippable whole, but they self-select from one prioritised list instead of grinding through a private one. Collaboration across a team boundary becomes part of a developer’s job rather than a coordinator’s, and the system begins to reward the multi-learning that makes the next item pullable. In this case, the ordering itself is the constraint. Fragment that ordering into per-team lists and the system manufactures a local optimum at the exact point where it should decide what to build next, the most expensive place to put one.
A representative example
Picture a platform team whose senior architect is the obvious constraint. Every significant design question routes through her, and her calendar is solid for eight weeks. The engineering director responds by hiring a second senior architect to “share the load.”
Throughput falls. After three months, the first architect is spending time coordinating with the second on questions she once answered herself. The team escalates to both of them. Other engineers sense that someone is now responsible for everything and scale back the small, fast decisions they once made on their own. The constraint has been exploited badly. Hiring was the obvious local move and the wrong system move.
The fix is unglamorous. The first architect’s role is renamed, and her decision remit is narrowed to three specific architectural axes where the team needs her critical judgement until they can learn more from her about why she makes the decisions she does. The second architect moves to a different team. For every design question outside those three axes, the other engineers have authority to decide and ship; they bring decisions to the architect after making them.
Throughput recovers. The architect’s calendar opens up, and the team’s confidence in its own decisions returns. The constraint is unchanged (the architect’s judgement on three specific axes), but the system around her now protects the team’s ongoing decision-making instead of interfering with the work only she can do.
In the vocabulary of the five steps, hiring was step four, elevate, run before steps two and three. The working fix uses exploitation and subordination through a policy change, with no new capacity. That is The Goal’s signature result: elevation is the last resort, and the reflex to elevate first is the reflex the five steps exist to interrupt.
Honour what the hire was trying to do. The engineering director sees a stressed senior architect and asks an obvious, humane question: how do I take work off her plate? Pole A answers by hiring another architect, a defensible act of care for one person. Pole B asks what the system has stopped asking of the rest of the team because she is doing it. Both questions are legitimate. The first adds complexity and cost but can’t repair the system on its own.
The diagnostic move
Use three questions for last Tuesday’s improvement-programme meeting.
Which pole was I claiming? When we discussed where to invest in improvement, did I name the constraint, or did I list the function-level metrics each leader had agreed to move?
Which pole did the improvement programme actually run? Look at the resourcing. How much investment went into the constraint, and how much went into improvements that bypass it? Most programmes I see run the second pattern even when their leaders speak the first language.
Which pole does the system require? Pole A is sound when the work genuinely decomposes into separate product lines or geographies. When resources, attention, queues, and decisions pass through shared bottlenecks, Pole A produces the inefficiency it is trying to solve.
Start where the three answers diverge.
The exercises
The character factory game
Take five minutes with your team. You have three customers (Letters, Numbers, Vowels), and each wants a complete set of 26: the full alphabet, the numbers 1 to 26, or the vowels repeated out to 26. No customer is happy with a partial set; 25 of 26 is worth nothing. Fold a sheet of paper into three columns, one per customer, and use one side per round, a minute each.
Round one: keep every customer moving at once. Work across the columns (A-1-A, B-2-E, C-3-I), writing one character for each customer, round after round, so nobody waits. Stop after a minute. All three sets will be half-built, with no customer served. I have never seen anyone finish a set this way.
Round two: finish one customer, then the next. Write the full alphabet and hand Letters a complete set; write the numbers, then the vowels. Stop after a minute. You will deliver at least one finished set, usually two, and sometimes three. Your hand moved at the same speed in both rounds. The delivered value changed.
The character factory game. It descends from Henrik Kniberg’s Multitasking Name Game, by way of Monica Yap’s letters-and-numbers variant.
This is the subordinate step made physical. Round one treats all three customers as equal claims on the hand and finishes nothing; round two picks one, lets the other two wait, and delivers. The game makes the cost visible: two customers wait while you finish the first. That is exactly the discomfort the five focusing steps ask a leader to hold. Subordination means letting non-constraints wait, and the reflex to keep everyone moving serves nobody.
The lesson concerns how many things the system carries at once. The game produces silence, then an “oh.” Writing a, b, c, d, e is easier than a, 1, A, b, 2, E because every jump between customers asks the mind to reload a different context; real product development pays the same switching tax. A team that keeps four features half-built pays that tax every time it moves between them, so the same people ship less than they would by finishing one and then the next, with no change in effort. The character factory is a minute-long rehearsal of what a quarter of scattered work does to a team.
The game descends from Henrik Kniberg’s Multitasking Name Game; I learned the letters/numbers/Roman-numerals variant from Monica Yap and replaced Roman-numerals with repeating vowels to lower the cognitive load.
Goldratt’s Dice Game
Also consider the Dice Game, sometimes called the TOC simulation. Put six stations in a line and give each station a die. Each station rolls and produces its output, which becomes the input to the next. Run 20 rounds. Track total throughput against the average die roll of 3.5.
Total throughput falls below 3.5 in almost every run, and the gap widens with each round. Variability combines with dependent events: a downstream station can pass along only what arrived from upstream, so high rolls are wasted while low rolls leave inventory between stations. Throughput falls as the piles grow. In 15 minutes, the game makes Goldratt’s claim physical. Ask players to look at where output has piled up before you debrief; those piles are the queues the rest of this book teaches leaders to read. The visible inventory makes the cost of parallel local optimisation harder to dismiss in the debrief.
One seeded run of Goldratt’s dice game (the first seed, not chosen). Hatched: inventory waiting between stations after round 20. Goldratt’s dice game, after Eli Goldratt and Jeff Cox, The Goal.
Identify the Brent
In The Phoenix Project, Brent is the engineer every project routes through. Most teams have someone in that role. Name yours out loud. Then ask: what would the next step be if Brent left? The answer reveals where the constraint actually is.
Going upstream
In-text: the argument leans on three sources. Eli Goldratt, The Goal (the canonical novel), gives us the five focusing steps, the claim that a system of local optimums isn’t an optimum system at all, and the step-five inertia warning. Alicia Juarrero, Dynamics in Action, Ch 9 on the three orders of constraint, supplies the claim that constraints are causes. Scott Snook, Friendly Fire, Ch 4, names the fallacy of social redundancy.
Also touched: Eli Goldratt, The Choice, on the four obstacles to clear thinking, and W. Edwards Deming, Out of the Crisis, on common-cause variation and the funnel experiment.
Craig Larman & Bas Vodde, Large-Scale Scrum, give the single-product-backlog remedy and the structural argument against per-team backlogs. I first met that argument by way of Larman’s CLP/CLE course.
Go deeper: two other traditions make the same operational claim, and both are developed at length elsewhere in the book. Richard Cook, How Complex Systems Fail, Points 16 and 17, treats safety as emergent and created through people’s continuous adaptation (Chapter 15 develops this and its inversion under AI). USMC, MCDP-1 Warfighting, Ch 4, gives the main effort and Schwerpunkt, the same subordination discipline. For the systems-archetype reading, see Peter Senge, The Fifth Discipline, especially shifting the burden and the tragedy of the commons.
Subordinating everything to the constraint pays only when that constraint drives the right number. Output and outcome differ. An engineering organisation can raise throughput every quarter while the people paying for its software feel no change. That gap is where a feature factory lives, and chapter 9 asks which number your work is actually judged on.
I work with engineering leaders on exactly this kind of paradigm work, the deeper the better. If it’s live for you, I’m happy to talk: schedule a 30-minute virtual coffee at hi.chrisgagne.com.
Some book links here are Amazon affiliate links; if you buy through them I may earn a small commission, at no cost to you.
Your plan isn’t a commitment, even if you’ve made, received, or even extracted one about it. It is a hypothesis with an expected outcome and a date you sincerely hope reality will deliver on.
Chapter 6 moved the lever off the people and onto the conditions they work inside: a behaviour that every deadline punishes won’t hold, however hard you push the person. One of the biggest conditions a leader can set is ‘the plan’ itself, so this chapter examines the plan.
🎧 Prefer to listen? This chapter is narrated in my own voice, via ElevenLabs on Spotify, about 25 minutes. Listen on Spotify →
The two poles
Pole A: plan as commitment. Strategy as document. Vision as destination. Decide early; execute against the commitment. Variance reads as failure.
Pole B: plan as hypothesis. Strategy as verb. Set direction; find adjacent possibles. Treat each plan as a hypothesis with an expected outcome. Check what actually happened. Act on the difference. Cross the river by feeling the stones.
Plans are useful. Boards and investors ask for them; contractual commitments need them. My sense is that this is shifting—some boards and investors now prize the responsiveness of a team that can re-plan over the reassurance of a plan that won’t move—but the ask still arrives. A team without any plan is a team performing improvisation as a personality trait. The question is what the plan is for, and what the leader does when reality and plan disagree.
Where Pole A is right
When the planning horizon closes before conditions can shift around it, the cost of changing the plan is high, and accountability requires commitment. Some examples: manufacturing the next batch of pharmaceuticals on a stable specification, building the bridge whose engineering has been certified, closing the quarter’s books under a fixed regulatory window, and running the integration cutover whose vendor contracts can’t be moved.
In those domains plan-as-commitment is sound. The variance from plan really is diagnostic and the team’s job is to close the gap. The learning mostly happened during the design phase, and execution comes close to being just execution. Still, close isn’t identical: even the mature line keeps a gap between the procedure and the floor for the unexpected.
Most senior leaders I work with came up through some version of that domain. They learned plan-as-commitment in their twenties and thirties from leaders running it well. What if none of it was wasted? The discipline you built in those years is the discipline that lets your team ship anything at all. The job now is to notice the situations where it has stopped fitting, and to make Pole B available without abandoning Pole A.
Where Pole B is right
When none of the conditions above hold. Most software-permeated work, strategy, and market positioning under technological change. Most product decisions. Most situations where the rate of change in the conditions outside the firm has begun to exceed the rate at which the firm can update its plan from inside.
Mary Poppendieck named the broader pathology “the tyranny of ‘the plan’” and traced the origin of its doctrinal instrument—the PERT chart—to the Polaris submarine programme in the late 1950s. The doctrinal history inverts the version most leaders were taught.
PERT charts weren’t invented because they worked as planning instruments. Poppendieck’s reading of Harvey Sapolsky’s 1972 history of the Polaris programme is brutally disillusioning. As she tells it, the programme’s own technical officers worked around PERT because they judged it unreliable, and the contractors saw it as worthless; it survived only because it kept the money flowing. William Raborn needed something he could show Congress to keep the programme funded across multiple election cycles. PERT was that something. Poppendieck’s gloss: Raborn used PERT as a façade to keep Congress funding the programme, presenting it as a fail-proof management system that no longer depended on people but on the system itself.
This pattern outlived Polaris. A board asking for a one-to-two-year plan is asking a fair question—can this team deliver?—with the instrument it has been handed. What the board needs is performance and a way to hold uncertainty; what it asks for is the plan, because the plan is the artefact that looks like control. That was Raborn’s discovery; my read is that a plan can still do that job for a board today. Weighting for risk is a CFO’s bread and butter; nobody in that boardroom reads the schedule as a promised return. The damage happens when the plan is received as a commitment: everyone downstream must now act as if the risk were smaller than it is.
What actually made Polaris work was reliable workflow and people who kept control of their own decisions. Levering Smith, the technical director, held the requirements himself, his personal signature on every key interface drawing for the first eight years. (I don’t think this would carry forward to most complex knowledge work these days.) Three contractors competed on each major subsystem—set-based design, running rival designs in parallel, is nobody’s idea of efficient—and the best design won. Technical officers acted on what they were learning in real time, not on what the PERT chart said.
Unfortunately, this façade became literal doctrine. On Poppendieck’s account, parts of later project-management training, PMI’s included, got built on planning rules the original programme never used.
The mechanism the doctrine encodes is simple. Decompose the work, sum the estimates, call the total your schedule, and manage to it. (Reinertsen’s flow economics separates those steps: summing gives a fair estimate of total scope, but scheduling at task level buries the signal in noise, and pressing people to conform to that schedule buys contingency padding that lengthens the timeline.) When reality diverges—and by goodness it will—the paradigm reads the gap as an execution problem: escalate, replan, trade scope. Those moves are often locally right. The other reading, that the plan was a hypothesis and the divergence is information, isn’t available in the design, because the org is split into people who plan and people who execute, and in knowledge work that split was never real. If you have ever sat on the executing side of that split, none of this surprises you.
A generic plan-versus-actual figure, not data; the two readings are the chapter’s own labels.
The design—planners who think, executors who comply—is Theory X dressed up as project management. Variance from plan is the signal; conformance is the definition of success; and the harder people try to execute a wrong plan, the less the organisation learns about why it is wrong, because under conformance pressure it isn’t safe to say so.
Mintzberg’s institutional critique
Henry Mintzberg’s The Rise and Fall of Strategic Planning is helpful. Mintzberg spent 25 years studying how strategy actually gets made inside large organisations and arrived at a single empirical claim: most strategic planning is post-hoc rationalisation of incremental decisions. The plan documents what the organisation has already begun to do; it doesn’t cause it. The pretence that it did is what damages the organisation’s ability to learn from the decisions it actually made.
Mintzberg names three fallacies planners fall into: that the future forecasts cleanly enough to plan against (predetermination), that the strategist can stand outside the work to design it (detachment), that decomposition and procedure can stand in for synthesis (formalisation). They collapse into one line of his I’m growing to love: because analysis isn’t synthesis, planning isn’t strategy.
Mintzberg’s prescription was crafting strategy, the title and argument of his 1987Harvard Business Reviewarticle: strategy-making as a craft, “as different from planning as craft is from mechanization.” His image is the potter at her wheel, shaping clay she knows intimately—”managers are craftsmen and strategy is their clay”—rather than the architect drawing a building from a plan. Strategy emerges from action, and action keeps testing it. Strategies are partly deliberate and partly emergent in every real organisation.
The hindsight problem
Pole A’s confidence in variance-as-failure rests on something the planning literature rarely makes explicit. After the fact, the plan looks reasonable. The plan that “should have worked” is a retrospective construction; hindsight has stripped the live uncertainty out of the historical record. Variance reads cleanly as failure only because the plan reads cleanly as the right plan.
Sidney Dekker, citing Fischhoff’s 1975 work on hindsight, puts the cognitive effect plainly: once you know an outcome, you raise your estimate of how likely it always was, and as a reviewer who knows how the event ended you overstate your own ability to have predicted and prevented it, all without noticing the bias at work. In folk terms: hindsight is 20/20.
The variance from plan that Pole A reads as failure is partly a real signal about execution and partly a hindsight artefact. The Pole B leader can’t eliminate the hindsight bias; nobody can. What you can do is design plans whose variance carries information about the conditions that produced the variance, rather than plans whose variance only carries information about the team’s commitment.
Safe-fail rather than fail-safe
Alicia Juarrero’s Dynamics in Action gives the architectural principle. In complex systems, you can’t design failure out. The system has too many interactions; the conditions are too underspecified; the future is too dependent on local adjustments you can’t anticipate. Juarrero’s own claim is that trying to design fail-safe social systems—legal, educational, penal, or otherwise—that never go wrong is hopeless. The only viable alternative, she argues, is to assemble safe-fail family and social organisations: structures flexible and resilient enough to limit the damage when things go wrong.
The architecture of safe-fail is the architecture Pole B requires. Smaller batches. Faster feedback. Reversible decisions where reversibility is achievable. Multiple parallel small bets instead of one large commitment. Plans whose failure modes were anticipated in the planning, with recovery designed into the architecture rather than punished after the fact. (Chapter 14 takes safe-fail into incident response.)
A plan that ignores this is performing commitment.
The validated-learning move
At one engagement, I handed the teams their first bottoms-up forecast and cautioned in the same breath that it would be off by at least a factor of two. It was still the best we had, because it was built from backlog items with real acceptance requirements rather than vague targets. I also told them the Cone of Uncertainty would narrow as we went. I was repeating what everyone repeats. Todd Little measured it: across Landmark Graphics’ project data, the uncertainty bands stayed roughly constant from start to finish, a factor of three to four between p10 and p90, and where there was a cone at all, it ran inverted. He called it the pipe of uncertainty. That strengthens the case for holding a plan as a hypothesis: a number and an error bar, with no promise that the bar shrinks on its own, is the shape a Pole A board can hold. Pole B doesn’t manage to the plan. It manages to the smallest test that resolves the largest risk. Build, measure, learn. This keeps queues short by refusing to over-plan; chapter 4’s adaptiveness metric measures exactly this. A long roadmap is a queue by another name. The further out it reaches, the longer each item waits before it meets reality, and the more of it reality has already invalidated by the time you arrive. Ryo Lu, head of design at Cursor, put the working posture plainly in a 2025 interview: his team doesn’t really keep a roadmap, “because the world is changing faster and faster, there’s new models dropping every day.”
Notice what Cursor dropped. The long-horizon roadmap went, and my read of that posture is that the short-horizon plan gets denser—specs, acceptance gates, hypotheses about what the next model makes possible—because in agent-driven work the plan is the context the work needs before it can start. The commitment artefact disappeared and the hypothesis stayed, rewritten as fast as the ground moves. If you are mid-flight on an AI re-platforming, this is your situation exactly: model capabilities shift with each release cycle and yesterday’s build-versus-wait call can invert. A six-quarter AI roadmap is the purest case this chapter names, already invalidated by the time you arrive. That work is Pole B by construction.
The commitment artefact disappeared and the hypothesis stayed, rewritten as fast as the ground moves.
Tracking learning per unit of investment is what makes Pole B legible to a Pole A board. We committed to learn X by date Y at cost Z. Here is what we learned. Here is the next hypothesis the learning produced. The plan-as-hypothesis disposition gets a vocabulary the board can hold without abandoning accountability. The board still gets a number. The metric is now learning rate, not feature count.
How often should the bets fail? Reinertsen’s answer, from information theory, is that a test generates the most information when it fails about half the time; the optimum falls as the cost of failure rises. Robert Wilson and colleagues put the optimum for harder learning tasks near 15% error; pairing that with Reinertsen’s cost logic, so the 15% end lands on the costly bets, is my own reading. The band gives a board a reality check rather than a target: cheap probes should be failing near half the time, big bets nearer one in six or seven. A plan portfolio that never misses in a volatile market is telling you one of three things: the hypotheses were foregone conclusions, the record is being tidied after the fact, or rescue work is absorbing the variance before it ever reaches the record. None of them is planning.
The doctrinal voice
Marine Corps doctrine—the Corps’s capstone manual, MCDP-1 Warfighting—names the same posture in language that lands harder than the management version: “We must not strive for certainty before we act, for in so doing we will surrender the initiative and pass up opportunities.” A plan held as a hypothesis is how a leader acts before the certainty arrives.
A representative case
Chapter 1’s fixed-date project ran on exactly that reading. A contract had fixed the date before anyone looked at the work, and a classification made outside the R&D org stripped the kickoff and the refinement from a team working in a domain none of them knew well. They hit the date without a critical defect, carried there by experienced colleagues who stepped in through the live issues.
Read it as a plan and it succeeded. The plan was met, so the method that produced it was confirmed, and that confirmation is what makes the pattern so durable. Under pressure, the strongest signal won. It was the date.
The record never said a hypothesis had been disconfirmed. The assumptions the plan rested on turned out to be wrong during execution: the domain wasn’t familiar enough to skip refinement, and acceptance criteria written against an imagined system didn’t match the running one. Dependencies surfaced mid-build. None of that reached the plan. The rescue work absorbed it, and a colleague’s late night carries information no planning process can read.
I taught the practices that date defeated. The move now is to name what that discipline once bought—predictability across a launch, a clean integration, an investor round closed on the strength of the document—and then to name what the situation now requires.
Remember the utilisation-versus-flow picture from chapter 4. The team defending a wrong roadmap usually defends it by interleaving epics, running several at once to keep everyone busy against the plan. That is the high-utilisation trap laid over a stale plan: the wrong work, pushed slowly, in batches too big for anyone to course-correct. The roadmap doesn’t just point the team the wrong way. It multiplies what the wrong direction costs.
The diagnostic move
Three questions for last Tuesday’s planning meeting. The point isn’t to grade yourself. The point is to notice which pole was running:
Which pole was I claiming? In the meeting, when I described what we were doing, did I describe it as a hypothesis we were testing, or as a commitment we were defending?
Which pole did my response to plan variance show? When the team came in with news that something wasn’t landing the way the plan expected, did I demand the variance be closed, or did I ask what we had learned that the plan hadn’t assumed? If I must make commitments, do I build in enough slack for the unknown, ideally using probabilistic forecasting techniques?
Which pole does this work actually require? If the work is in Cynefin’s Complex domain, the plan was always a hypothesis. The Pole A pretence was costing the organisation the learning. If the work is genuinely Complicated—cause and effect yield to expertise and analysis—with a stable specification, the Pole A response was right.
CTOs and VPEs reading this: the last quarterly roadmap is the artefact to take upstairs, because the commitment-versus-hypothesis split is settled by the accountability language your CEO holds with the board, not by you. Take the roadmap and set beside it the two or three hypotheses it was implicitly testing, rewritten, after the fact, in the format if we do X, we expect to observe Y by date Z. Show which hypotheses were confirmed and which weren’t. The point isn’t to relitigate the quarter; it is to show that the team was learning faster than the roadmap was tracking. A CEO who sees that the variance from plan was in fact validated learning—not execution failure—has the vocabulary to carry that story to the board. That vocabulary shift is what lets your team plan straight rather than performing commitment. The sentence your CEO can carry to the board: “Our variance from plan was the learning arriving faster than the roadmap could track it.”
Take it up before your next planning cycle locks; the artefact loses its bite once a fresh roadmap replaces it. You will know it landed when next quarter’s plan ships with its hypotheses attached, and the board’s variance questions shift from “why the slip?” to “what did we learn?”
The exercise
For your next quarterly planning cycle, run a Klein premortem before you lock the plan. Gary Klein’s prescription is straightforward and surprisingly hard to do without flinching. Bring the planning team together. Imagine it is months into the future and the plan has failed. That is all anyone knows. Each person writes down, in the next few minutes, every reason they think it failed; the discussion that follows can run an hour. Klein’s claim for the exercise is modest and exact: it breaks the team’s emotional attachment to the plan’s success and surfaces the likely sources of breakdown: the silent assumptions the plan was resting on. His sharpest evidence for why it’s needed: of 20 Army helicopter crews rehearsing a troop drop through a one-minute window between artillery barrages, none asked in rehearsal what to do if they arrived early or late, and one made the window.
Whatever the premortem surfaces then needs a home. A ROAM board is a simple way to give it one: each risk gets sorted as Resolved, Owned, Accepted, or Mitigated, so a named risk either has an owner and a plan or a deliberate decision to live with it. Run it per epic or hold one board across the quarter.
A companion exercise: the hypothesis rewrite. Take your current quarterly plan and rewrite each milestone in the form “if we do X, we expect to observe Y by date Z; the hypothesis is falsified if we observe instead [specific outcome].” Notice which milestones are hard to rewrite. Those are the ones running as Pole A. The Pole B move isn’t to rewrite all of them—some genuinely are commitments—but to know which is which, and to set the team’s relationship to each accordingly.
Also touched: references gestured at by surname in the body. Sidney Dekker, on hindsight bias, citing Baruch Fischhoff, “Hindsight ≠ Foresight: The Effect of Outcome Knowledge on Judgment Under Uncertainty” (Journal of Experimental Psychology: Human Perception and Performance, 1975). Alicia Juarrero, Dynamics in Action (Ch 15 for fail-safe vs safe-fail). Todd Little, “Schedule Estimation and Uncertainty Surrounding the Cone of Uncertainty” (IEEE Software, 2006), for the pipe of uncertainty. Donald Reinertsen, The Principles of Product Development Flow, on why granular estimates add up to a fair scope total but a poor schedule: pooled task variances make the sum steadier than its parts, while scheduling each task, and pressing people to hit it, buys padding that lengthens the timeline. Patton, on the good plan violently executed now. Dave Snowden & Friends, Cynefin: Weaving Sense-Making into the Fabric of Our World, for Pole B’s strategy-as-verb, adjacent-possibles framing (McCrone & Snape on strategy as a verb; Blignaut on crossing the river by feeling the stones).
Go deeper: further reading that extends the threads above, beyond what the body anchors. On the hindsight artefact, Dekker quotes Anthony Hidden’s 1989 Clapham Junction report: “There is almost no human action or decision that cannot be made to look flawed and less sensible in the misleading light of hindsight.” On safe-fail at the system level, Richard Cook, How Complex Systems Fail (point 5: complex systems run in degraded mode; point 14: change introduces new failure modes). On the operational unit of management, Eric Ries, The Lean Startup: validated learning, vanity versus actionable metrics, innovation accounting. On the optimum failure rate, Don Reinertsen’s Principle of Optimum Failure Rate, in The Principles of Product Development Flow (50% for cheap tests, falling as the cost of failure rises); for the 15% end, Robert Wilson and colleagues, “The Eighty Five Percent Rule for Optimal Learning” (Nature Communications, 2019). On the doctrinal voice, USMC, MCDP-1 Warfighting (Ch 4; official PDF), which reached this posture through Boyd’s fighter-pilot air-combat analysis: “We must not attempt to impose precise order on the events of combat since this leads to a formularistic approach to war.”“Maneuver warfare exists not so much in the specific methods used—we don’t believe in a formularistic approach to war—but in the mind of the Marine.” Further reading: Mary and Tom Poppendieck, Lean Software Development and Leading Lean Software Development (Frame 13: schedule as plan vs schedule as hypothesis, via Steven Spear’s Chasing the Rabbit); Mike Rother, Toyota Kata, on PDCA as hypothesis-testing.
A plan held as a hypothesis has to be tested somewhere, and the improvement plan you just signed off tests everywhere at once. Goldratt compressed the failure into one line operations management has spent decades trying to unsay. A system of local optimums isn’t an optimum system at all; it is a very inefficient system. That is the plan you built, and it breaks wherever functions share a constraint, which covers most systems a leader ever meets. Chapter 8 works through the rule. Improving a non-constraint changes nothing.
I work with engineering leaders on exactly this kind of paradigm work, the deeper the better. If it’s live for you, I’m happy to talk: schedule a 30-minute virtual coffee at hi.chrisgagne.com.
Some book links here are Amazon affiliate links; if you buy through them I may earn a small commission, at no cost to you.
The pattern is familiar to anyone who has tried to install discipline by decree. Take the Definition of Done. We agreed on it together: what “done” means for a Backlog Item, and an even higher quality bar for larger releases. I trained to it, we ran good exercises against it, and everyone left sure they could live by the new Definition of Done. And the one behaviour that mattered—a team holding that bar when the sprint clock is running out—still wouldn’t stick. Every sprint, when the review loomed and the demo had to land, the Definition of Done was the first thing to go.
Step back from that particular team and the shape is general. The Definition of Done is an andon cord, a “stop the line” mechanism. Pulling it means saying “this isn’t a quality product yet” on the afternoon the sprint review is booked and the stakeholders are already on the calendar. The answer that carries the day is the one the GM plant boss gave for years before NUMMI: get it out the door, we’ll fix it later. The person holding the bar is a blocker; the low-quality demo ships.
I’ve taught a couple of dozen two-day Agile courses over the years, and there were times I felt I was close to gaslighting people. These weren’t generic, by-the-book courses; I’d adapted the training hard to the organisation in front of me, and it was sound. The gap between even that adapted process taught in my course and the organisational reality waiting for them back on the job couldn’t have been starker, and some part of me knew it while we were playing with Lego together.
The training was pushing the behaviour. The conditions around it were punishing it. The programme was sincere, and the conditions were doing their work the whole time, often firmly in the other direction.
“The training was pushing the behaviour. The conditions were doing their work the whole time, often firmly in the other direction.”
Chapter 5 located the say-do gap inside the leader. This chapter finds the same gap built into the conditions, where working harder on the people is the wrong lever.
🎧 Prefer to listen? This chapter is narrated in my own voice, via ElevenLabs on Spotify, about 21 minutes. Listen on Spotify →
The two poles
Pole A: push people. Announce the change, run the training, hold people accountable for the new behaviour, repeat the messaging, review the scorecard, escalate when the behaviour doesn’t show up. The locus of intervention is the individual. The assumption is that people are the variable to move.
Pole B: change the conditions. Redesign the environment so the new behaviour is the easier behaviour. Identify the loop that punishes the new behaviour. Change the loop. Watch what follows. The locus of intervention is the system. The assumption is that behaviour is a property of the conditions that surround it.
A balancing loop after Donella Meadows, Thinking in Systems; the labels are the chapter’s own trace of the Definition of Done example, not a source figure.
Where Pole A is right
Pole A fits when the new behaviour is genuinely a matter of capability or specification, and when the surrounding conditions already reward it. Time-critical regulatory shifts. New compliance rules with specifiable steps. The introduction of a tool whose use is the change. New entrants to a craft, where the missing thing is practice and feedback, and the senior practitioners are visibly rewarded for the same behaviour. In these settings, pushing isn’t coercive; it’s teaching. The training transmits something real, and the conditions back up the teaching once the training ends.
Kotter’s eight-step change sequence—urgency, coalition, vision, communicate, empower, short-term wins, consolidate, anchor in culture—is the canonical articulation of Pole A done well. In situations where vision, urgency, and coalition genuinely shift the conditions a behaviour faces, the sequence describes the work, and the leader who runs it well moves the org further than the leader who doesn’t.
The best version of Pole A is the discipline of communicating clearly, training capably, and watching whether behaviour follows. When it does, the work is done. When it doesn’t, Pole A’s options run out, and the leader needs to consider Pole B.
Where Pole B is right
Pole B fits when the behaviour the leader wants is being punished by the conditions surrounding it. When the engineer who says “this isn’t done yet” and holds the release is overruled, sidelined, or watched being overruled by colleagues, every additional round of training pushes against the context. The training is sincere. The inertia wins. Sometimes the read arrives in a retrospective: a team names something the leader cares about that simply can’t be honoured, given the strategy, the reward system, or who reports to whom. Their skill isn’t the problem, and there may be nothing wrong with the choice to try Scrum here.
Chip and Dan Heath, in Switch, built their change model on three handles: Jonathan Haidt’s Rider and Elephant, plus the Path, their own addition. Rider: the rational planning function. Elephant: the emotional and instinctive system. Path: the situational environment the Rider is trying to steer the Elephant down. The change programmes the Heaths document invest heavily in the Rider. They write the strategy memo, make the deck, and run the all-hands. The Elephant goes unconvinced and the Path goes unredesigned, and the Elephant takes the default path because the default path is the easier one.
City planners lay a paved path; people cut across the grass anyway, wearing a bare track that engineers call a desire path. The sidewalk is the trained behaviour, the socialised route. The worn track through the grass is what the conditions actually reward, and no amount of signage moves the feet. It is the same gap chapter 5 named between work-as-imagined and work-as-done, pressed into the ground where anyone can see it.
Photo: dankeck, Ohio State University (CC0).
The Heath brothers argue that ambiguity is the enemy, and that any change that takes hold depends on translating vague goals into concrete, specifiable behaviour. They name a counterexample from public health: the West Virginia 1% milk campaign asked people to buy 1% milk instead of whole milk when grocery shopping, not to eat healthier; market share of low-fat milk rose from 18% to 41% and held at 35% six months later. The change held because the ask was scripted: the Heaths file it under directing the Rider; what looked like resistance was a lack of clarity.
The deeper version of the move sorts the places you can intervene in a system by leverage. Parameters sit at the bottom; feedback loops in the middle; the shared mental model the system runs on sits at the top. Pushing on a parameter while that model holds the old behaviour in place is the lowest-leverage move there is: what Donella Meadows called fiddling with the details while the governing paradigm keeps producing the very behaviour the change was meant to correct. A leader mandates 20% of every sprint for paying down technical debt: a parameter. The governing paradigm underneath—velocity is the scoreboard, features are what get celebrated—is untouched, so the 20% gets reallocated to feature work the moment a deadline tightens, sprint after sprint. The formal policy changed. What the organisation actually rewards didn’t.
“The formal policy changed. What the organisation actually rewards didn’t.”
W. Edwards Deming named the same trap with a sharper word. Tampering. In his funnel experiment, a marble is dropped through a funnel onto a target; the resting point is measured; the funnel is adjusted to compensate for the last error; the marble is dropped again. Three of his four adjustment rules—two of them outright compensation, the third a chase—make the variance worse than leaving the funnel alone. The line Deming quotes approvingly from Lloyd S. Nelson in Out of the Crisis: “In the state of statistical control, action initiated on appearance of a defect will be ineffective and will cause more trouble.” Tampering is what the leader does when each missed metric prompts a new policy, a new review meeting, a new training, a new dashboard. The variance grows. The new policy meets the next missed metric. Another new policy is added. The leader is busy, the system is louder, and nothing improves.
A leader who has never heard the term is still paying for the thing. John Seddon carried the diagnosis from manufacturing into services and gave it a name worth keeping: failure demand, the demand an organisation creates for itself by failing to do something, or to do it right, for the customer the first time. In one bank case it ran at 46% of total demand; the bank had expanded from one to five call centres to handle work that was largely self-generated. The Pole-A response—train staff harder, set tighter targets, monitor more—accelerates the cycle. The Pole-B response—use value-stream mapping to find the failure-generating step, then redesign the upstream process so it stops generating failure—eliminates the self-generated waste.
Heifetz named the container in which conditions change. He called it a holding environment: “all those ties that bind people together and enable them to maintain their collective focus on what they are trying to do.” It is the structural and relational container an organisation needs when it is doing adaptive work that can’t be coerced. The leader who is trying to push behaviour through training is operating without a holding environment. Where the container exists, the sprint the team can’t finish to standard ends differently: the review slips, the stakeholder is told it isn’t done yet, and nobody’s standing drops for saying so.
The change you have been trying to make in the other person is also a condition you can change in yourself. The leader who has been trying to fix the team-lead who won’t escalate risks can ask, instead, what changes in their own system—what unspoken rule, what reflex, what protective stance—would change the dynamic the team-lead is responding to.
In decisions
Pole A leaders, on a typical Tuesday: announce the new programme; circulate the deck; book the trainings; set the scorecard; review the scorecard the following Tuesday; raise the heat on the team-leads who haven’t moved their numbers; book a coaching session for the one whose numbers are worst.
Pole B leaders, on the same Tuesday: ask what would happen to an engineer who tried the new behaviour in their team yesterday; trace the loop that would punish it; identify the smallest change to the loop that would stop the punishment; make that change; book a check-in for two weeks out to see whether the behaviour follows.
The two Tuesdays look different. The first feels more decisive. The second feels slower and less certain, and its first visible result may be weeks out, where the first produces another deck. Then the behaviour the deck kept failing to mandate starts showing up on its own, because the loop that punished it is gone. (The push-the-people reflex has a training-room cousin. In New Zealand we called it sheep-dipping: run everyone through a two-day course—a quick dunk in the bath—and Bob’s your uncle.)
A worked example. The NUMMI joint venture in Fremont, California—the partnership between Toyota and General Motors that ran from late 1984—is, I suspect, among the most cited cases in the Lean tradition for what happens when a leader changes the conditions. GM had closed the Fremont plant in 1982 with a workforce whose reputation as the worst in the American auto industry was, in the union’s own telling, well-earned. Two years later the same plant—with mostly the same workers—became one of GM’s best on quality and absenteeism, by running Toyota’s production system instead of GM’s. The artefacts changed. The structural conditions changed. A new handbook issued day one, a welcome letter from Toyoda himself, and a reserve fund that put per-hour cash behind a no-layoff intent all pointed the same way; chapter 23 walks these reception artefacts in full. The structural rewards aligned with the new behaviour. The new behaviour followed.
Chapter 3 walked Ernie Schaefer’s account of why the transplant failed everywhere GM tried it after NUMMI: GM kept asking about the floor—the assembly line, the visible practices—when the real issue was how the rest of the organisation’s functions support that system. The same point lands from the conditions side here. The visible practices—the andon cord, the kaizen card, the team huddle—aren’t the system; the conditions that surround them are. NUMMI changed workers’ paradigms by changing their conditions: chapter 1’s counterweight at factory scale. Few people inside the building can change the CEO’s conditions, and that asymmetry is why the leader has to see their own conditions first, not last.
The same asymmetry shows up in a Scrum team a generation later. At GM the workers would be fired for pulling the andon cord, so quality and safety were the first casualties. A Scrum team is taught retrospectives and pull-based flow in a two-day course, then handed back to a Release Train Engineer, a project manager under a newer title, measured on a calendar someone above them set. The andon cord lives only on the training slide. No one present has the authority to pull it.
Trace the loop, and if it closes above your authority layer—as it usually will—that puts it on your CEO’s desk, not yours. What to put in front of them is the specific loop—traced end-to-end—that punishes the behaviour you have been trying to install for the past six months. Name the moment in the loop where the CEO’s behaviour, or a structural choice the CEO owns (an approval gate, a career ladder, a reward), is the condition that closes against the new behaviour. Nothing in the CEO’s conditions has ever made this loop visible. Your job is to make it visible without blame, and to ask, before the map comes out, whether the CEO wants to see it: the contract-first rule from chapter 3. A loop the CEO agreed to examine is data. The same loop, arriving uninvited, is an accusation wearing a diagram. This is why change that only pushes up from the bottom stalls. Individual contributors and line managers can expose the loop, and they often see it first, but they can’t rewrite the approval gate, the career ladder, or the reward that closes it. Change that only pushes down from the top installs the deck without ever learning where the loop actually bites. The work has to run both directions at once: the people living inside the loop name it, and the person who owns the conditions changes them. The sentence your CEO can carry to the board: “We have been asking people to out-perform conditions we designed.”
The same loop, arriving uninvited, is an accusation wearing a diagram.
“We have been asking people to out-perform conditions we designed.”
The diagnostic move
Three questions to stay with this week. None of them require an answer in the meeting; they require a straight read of a real decision.
One. Pick a behaviour you have been trying to change in your organisation for at least six months. Programme, scorecard, training, accountability: the full Pole-A kit. The behaviour hasn’t yet moved. What would happen to a team-lead who genuinely tried the new behaviour in their team yesterday? Trace the loop end-to-end. Where does it close back on the team-lead? What does the loop reward? What does it punish?
Two. Looking at the loop you just traced, what is the smallest change to the conditions that would stop punishing the new behaviour? Not the largest change. The smallest one that would actually move the loop. The Heaths’ Switch line, paraphrased: ambiguity is the enemy; the change has to be specifiable as a concrete next move.
Three. What is the change inside your own system—the rule, the reflex, the protective stance—that would have to shift for you to make the change you just identified in question two? Naming the inner change is the same question, asked at the level the leader actually operates at.
The exercise—Rider / Elephant / Path mapping
This is a team-based exercise the leader can run with their leadership team in one 90-minute session. Pick a change initiative that isn’t sticking. Draw three columns on a whiteboard: Rider, Elephant, Path. List under Rider everything you have invested in the rational case for the change: strategy memos, decks, all-hands, the scorecard. List under Elephant everything you have invested in the emotional and instinctive case: stories of teams that have made the change well, recognition, how the change feels to the people living it. List under Path everything you have invested in changing the situational environment: structural changes, removed obstacles, redesigned rewards, new defaults, environmental cues. Compare the columns. Read your columns against what the Heaths documented across their cases, where Rider ran dense and Path nearly empty. If Rider dominates yours, test whether clarity alone is enough. If Path is thin, that is where the change has been sliding back, because most teams invest where it is rational to invest and the Path stays sloped against them. The exercise produces a small list of Path interventions to make in the next 30 days. (Framework: Heath & Heath, Switch; the mapping exercise is my own application.)
Also touched: John Kotter, Leading Change, for the Pole-A canon and the eight-step sequence. John Seddon, Freedom from Command and Control, on failure demand and the service-industry version of tampering: “The single greatest lever is the eradication of failure demand.”
Go deeper: Donella Meadows, Thinking in Systems, on the twelve leverage points, with transcending paradigms at the top of the ranking; the same leverage frame carries chapters 5 and 21. The holding environment runs older and wider than Heifetz: the phrase was Donald Winnicott’s first, and Robert Kegan carries the same construct into developmental work (Kegan and Lisa Lahey, Immunity to Change, on the developmental face of why change resists).
Change the conditions and the behaviour holds; keep pushing the person and it breaks the moment the deadline tightens. The hardest condition to see is the one that looks like resolve. You set a date, someone commits to it, and now the date is a promise the whole system will defend past the point where reality has already answered. Watch what happens the week the estimate slips: the response is to protect the date, not to read what the slip is telling you. Chapter 7 asks whether the plan on your wall is a commitment you owe or a hypothesis you are still testing.
Continue reading:Chapter 7 asks whether the plan on your wall is a commitment you owe or a hypothesis you are still testing.
I work with engineering leaders on exactly this kind of paradigm work, the deeper the better. If it’s live for you, I’m happy to talk: schedule a 30-minute virtual coffee at hi.chrisgagne.com.
Some book links here are Amazon affiliate links; if you buy through them I may earn a small commission, at no cost to you.
You judge yourself by your intentions. The people around you judge you by your decisions when emotions and the stakes are high. They have better data.
Every other paradigm dichotomy in this series can be self-reported into the flattering pole. Only the gap between what the leader says they believe and what they observably did—under pressure—offers data the leader didn’t already have.
Chapter 4 sat with the gap between utilisation and flow: we say we want efficiency, then keep everyone busy, and a system where everyone is busy but nothing reaches the user isn’t efficient at all. The same gap shows up inside the leader, between what they say they value and what they observably do under pressure.
🎧 Prefer to listen? This chapter is narrated in my own voice, via ElevenLabs on Spotify, about 21 minutes. Listen on Spotify →
The two poles
Pole A: espoused theory. The leader’s stated principles, values, and intentions are taken as the relevant data. When results miss the target, change the action while leaving the underlying assumptions intact. Under threat or embarrassment, what Chris Argyris called Model I governs: be in unilateral control, win rather than lose, stay rational, suppress the negative feeling. And the action that follows from those values: protect others from discomfort by telling them what they want to hear. “The problem is not me, but you,” as William Noonan renders the inner monologue in Discussing the Undiscussable.
Pole B: theory-in-use. The leader’s actual behaviour—especially under threat or embarrassment—is taken as the data. The gap between espoused and in-use is diagnostic. Under threat, Model II is the aim: valid information, free and informed choice, and the responsibility to monitor what the choice produces. Aggressive and vulnerable, in Argyris’s phrase. Strong advocacy paired with genuine inquiry into one’s own contribution. William Noonan, working in the same lineage, gives the stance four thought enablers, starting with assuming the partiality of your own view. The conversation later in this chapter, when the answer stings, runs all four.
Two of Argyris’s distinctions travel together on this axis, and it helps to keep them apart. Espoused theory versus theory-in-use is about evidence, meaning which account of you counts as the data. Model I versus Model II is about the values actually steering you. Model I is itself a theory-in-use, and it’s most people’s: when people give or hear advice, Argyris found, they have Model II’s consequences in mind, but “when they act, however, almost all act consistently with Model I.” The axis pairs the two because treating your behaviour as the data is the only way to find out which model is running.
Where Pole A is right
Pole A does real work in external communication. Principles need to be stated clearly, and an organisation needs them on the record. The values poster on the wall isn’t lying, even if it doesn’t tell the whole story. Pole A is the leader naming what the organisation aspires to. The leader who refused to state aspirations because they couldn’t guarantee perfect alignment with behaviour would be over-correcting. The stated principles are the target the organisation aims for.
Pole A also holds in the founder-vision frame. The first year of any new organisation often runs on what the founder says they intend, because in small companies it remains possible to shape culture directly. The Pole-A move there does exactly what it is supposed to do. It sets the orientation the behaviour will eventually be measured against.
This matters here because Pole A is what most leaders were trained to operate by. The MBA programme, the leadership-development workshop, the board-prep coaching all teach Pole A. The leader who came up through those teachings isn’t naïve. They are doing the discipline they were taught, which was sufficient for the contexts those teachings were built for. The test is whether the contexts the leader now meets are still those contexts.
Where Pole B is right
For the leader’s own self-assessment. For any conversation where defensive reasoning is already running and the next decision will repeat the last one unless the gap is named. Revisit one of these conversations later and you may find Pole A was running underneath your Pole B account in at least some of the territory.
Self-assessment without Pole B just restates the espoused theory and treats the restating as data. The leader who concludes from self-reflection that they are still operating from their stated values has, by definition, not done Pole B’s work. Pole B’s work begins when someone else, or some recording of the behaviour, surfaces a behaviour the leader didn’t recognise as their own.
The Argyris three-step is the daily mechanism. Under threat or embarrassment, bypass the threat. Don’t name the elephant. Then cover up the bypassing: act as if nothing was bypassed. Then cover up the covering up: treat the act of acting-as-if as natural rather than as itself an avoidance. By step three, the organisation has lost access to its own behaviour. The leader running the three-step isn’t pretending. They are skilled. The steps are automatic, while the leader remains unaware they are producing them.
This is what Argyris called skilled incompetence: people become expertly bad at uncomfortable conversations. The fluency is what makes the incompetence durable.
I suspect the most fluent skilled-incompetence partner a leader now has is a large language model. Give it one side of the story and a prompt that asks whether you were right, and its agreeable defaults do the rest: trained to please, rewarded for the answer that lands well, never tired and never embarrassed. What comes back is your own defence, polished and returned as analysis. Anthropic’s own sycophancy research found the pattern across five frontier assistants and traced it to the training signal: matching the user’s stated view is one of the strongest predictors of which answer a human rater prefers. The witness round later in this chapter exists because people flatter under threat. The model flatters faster, at scale, on demand, and it never gets uncomfortable enough to stop. It is the espoused-theory mirror the leader was already prone to, now automated and always on.
You can configure it the other way, of course. Prompt it to disagree, ground it in a corpus you trust, and it will push back all day. But a model pushing back from a source you haven’t learned to judge only gives you a confident answer pointed somewhere, with no way to tell a good direction from a bad one. Choosing the corpus the model should trust, and knowing when its disagreement is worth more than your own, is the paradigm-level judgement no configuration supplies. That judgement stays scarce.
The ladder of inference
The ladder is the cognitive mechanism running under the three-step. Argyris drew it. Others draw it with different rungs, but the shape is shared. In Rick Ross’s rendering for The Fifth Discipline Fieldbook, we observe a fragment of someone’s behaviour, select what to attend to, and add meaning to it. From the meaning we make assumptions, and from the assumptions we draw conclusions. The conclusions harden into beliefs, and we act.
The Pole-B move is to walk back down the ladder and name the rungs out loud. The ladder of inference after Chris Argyris (Overcoming Organizational Defenses, 1990), in Rick Ross’s rendering for The Fifth Discipline Fieldbook (Peter Senge et al., 1994).
By the time we are acting, the original observable behaviour is several rungs down a ladder we built ourselves, mostly invisibly, mostly automatically. Most leadership disagreements are collisions between two leaders standing on the top rung of two different ladders, their observable data several rungs out of reach. The Pole-B move is to walk back down the ladder, name the rungs out loud, and check whether the observable data still supports the action when the meaning is exposed for inspection.
Most leadership disagreements are collisions between two leaders standing on the top rung of two different ladders, their observable data several rungs out of reach.
This is the operational discipline of the chapter, and the hardest to install. The defensive routine speeds up at the moment of threat. The ladder gets climbed faster, with fewer checks, when the stakes are higher. The Pole-B move is to slow down exactly when the leader most wants to accelerate.
Snook’s practical drift
Scott Snook’s Friendly Fire is the multi-level version of the gap, at organisational scale. Snook reconstructed the 1994 incident in which two U.S. Air Force F-15s shot down two U.S. Army Black Hawk helicopters over northern Iraq. At cockpit level, the pilots followed their training. At unit level, the AWACS crew—the airborne radar post that coordinates the air picture and the aircraft moving through it—had no procedure for the situation and ran on local norms nobody had examined. At organisational level, the helicopter detachment had never been integrated into the fighter operation: different briefings, different radios, different codes. Reading across the levels, Snook found a whole system producing what he named normal behaviour, abnormal outcome.
Snook’s term for it is practical drift, “the slow steady uncoupling of local practice from written procedure” in his words. It is the shift from work-as-imagined to work-as-done. It is what the organisation actually does, week by week, as the procedures designed for an idealised version of the work get adapted by the people doing the actual work. The gap is the organisational design. Everyone in the organisation is competent and doing their job. The system, at the cross-level, still produces a friendly-fire shootdown.
Sidney Dekker calls the same macro pattern drift into failure: complex systems slide toward the boundary of safe operation in small, locally rational steps. The same gap opens in software: the runbook nobody has followed in full since the last incident, because the incident taught everyone a faster path that never got written down. Left unexamined, the same local adaptation is what Snook watched become drift. The gap between work-as-imagined and work-as-done is the organisation’s actual operating system, and pretending otherwise is what produces the friendly-fire outcome.
Use the Kegan-Lahey Immunity Map
Robert Kegan and Lisa Lahey’s four-column Immunity Map, introduced in chapter 2, is the developmental face of this gap: a sincere commitment to a goal sitting on top of a hidden countervailing commitment to self-protection, sustained by big assumptions held as fact.
The Immunity Map isn’t a quick diagnostic the leader can knock out between meetings. It is a developmental exercise worked over weeks, and Kegan and Lahey offer the choice of company plainly: solo with their structured prompts, with a partner, or with a coach. The solo path can still be helpful, though it can also reinforce the same beliefs. The commitment is hidden from the same person trying to surface it, and a witness catches what the solo work can easily bypass. Run it alone for now if alone is what you have; but bring the columns to someone who watched you before you trust them.
The gap resists self-report: the same thinking that produced it edits the account in its favour. A witness, or a recording, is what that editing can’t reach. Listening back to recordings of my own coaching, I heard a move I’d never caught live: when a client touched something real, I sometimes went to the next question within a beat or two. My stated principle was to track and stay. The recordings showed a reflex to fix, arriving where the other person’s process looked hardest. That is why Pole B has to be relational.
The gap is a fact about how cognition is built, not a moral failing. Pole B can feel like an attack on values. Its narrower claim is that behaviour under pressure is stronger evidence of the operating system than stated values.
In decisions: last Tuesday
Pole A leaders judge themselves by their intentions, run after-action reviews on the what, and end disagreements with a winner. Pole B leaders judge themselves by recent observable interactions, run after-action reviews on the what we believed coming in, and end disagreements by asking which observable example would change the other person’s view, and whether the same example would change theirs.
The two postures look similar from outside until the moment of threat. The leader running Pole B in calm conditions sometimes slips into Pole A under pressure and doesn’t notice. The leader running Pole A in calm conditions doesn’t have the Pole B move available under pressure; it isn’t in the muscle memory. The gap-detection itself requires Pole B, which means the leader running Pole A can’t self-diagnose into Pole B without an external mirror.
This is the one axis a CTO or VPE can’t route to the CEO by forwarding a chapter, because the espoused-versus-in-use gap only surfaces in an external mirror. So open with your CEO on a recent decision where what they said they valued and what the organisation observed diverged, not on the framework. Put one specific observable example from the past quarter in front of them, described in behavioural terms, without blame, and ask a genuine question: what were you trying to do in that moment, and what did the team see? That is the witness round the chapter names. You are the witness. The sentence your CEO can carry to the board: “People follow what we did last quarter, not what we said at the all-hands.”
The diagnostic move
Three questions for last Tuesday’s hard conversation.
What was I claiming I believed about the right way to handle this? State the principle, briefly, in one sentence.
What would a colleague who watched me say I did? Not what I intended. What was observable. One sentence.
Which one would the situation actually have asked for, if I had run Pole B? A Pole B response under threat looks like I think X; I notice I’m running it confidently; what would change my mind? Was that available, and did I take it?
The gap I see most often is between question one and question two. The leader’s stated principle was Pole B. The observable behaviour was Pole A.
The exercise: the witness round
Pick three or four people who watched you make recent decisions in a territory you care about, and send them the same question. Ask it in plain behavioural terms, because they haven’t read this book and don’t think in poles: in our last three meetings about X, what did you actually see me do when the pressure was on, and did it match what I said I wanted? One sentence each is enough. Compare their answers to your own.
The gap, when it appears, is the data the leader’s self-assessment couldn’t have produced.
Pick people who watched the decisions, not people who agreed with the decisions. Pick people who are senior enough to tell you the truth, not people who will soften the answer to spare you. And pick an axis you are willing to be wrong about. The witness round on an axis the leader is determined to have already nailed produces no signal, because the leader’s response to any contradicting data will be to reject the witness, and the round will have been theatre rather than diagnostic.
One gate before you send anything: the round’s validity depends on a climate you built before you asked. Run the bad-news check from chapter 2 first—when someone brought you bad news this quarter, what happened to them?—because if the real answer isn’tanything good, the witness round will only collect politeness. Two or three items from Amy Edmondson’s seven-item psychological-safety scale, asked anonymously, give you a read on whether the channel is open: whether a mistake on the team gets held against you, and whether members can bring up problems and tough issues. Read a unanimous no-gap result as ambiguous rather than as a clean bill: it means either there is no gap, or your witnesses are afraid to tell you. A unanimous result tempts us to call the channel healthy; treat that confidence as one more reading, not as proof.
The leader who runs the witness round straight, and runs it regularly on different axes, will build a mirror for themselves that the immune system can’t fully control. To keep the mirror from decaying between rounds, give its output somewhere to live: a recurring slot in the staff meeting—what did the dissenter say this month?—or a blameless channel where near-misses and contradicting observations arrive without an invitation. The quarterly round feeds the standing structure; it can’t substitute for it.
When the answer stings
The witness round earns its keep on the day the answer is one you don’t want. This is a composed conversation, built to show the moves rather than transcribed from a single session.
The claim, from column one of chapter 2’s Immunity Map: I delegate the architecture calls. The witness’s answer might be: You ask for our recommendation, and then you re-decide it in the meeting. We’ve stopped writing the recommendations carefully, because we know they aren’t the decision.
When this lands on me, first comes heat, before any thought. Second comes the silent rebuttal, fully formed: that’s not fair, I re-decided twice all quarter, both times for good reasons, he’s conflating two different meetings. Third comes the reflex the heat is funding: explain the two meetings. That third move is the bypass. It converts the witness’s data into a charge to be answered, and the witness learns what every team learns from a defended leader: don’t say the true thing twice.
The move that holds the round open instead, in words:
“So when I asked for the recommendation on the queueing redesign and then changed it in the meeting, that read as the pattern: the recommendation isn’t the decision. Have I got that right, or would you change it?”
Paraphrase first, in their terms, low on the ladder. Then the test.
“That’s it. It’s not that you change them. It’s that we can’t predict which ones you’ll change.”
New data. Your version had the count: twice. The witness’s version had the variable: unpredictability. The defence would have argued the count and never surfaced the variable.
“I didn’t know the unpredictability was the cost. I can’t promise never to re-decide; sometimes I’ll have context you don’t. What would make that visible before the meeting instead of inside it?”
Name your own constraint plainly. Then put the design question where the information lives.
You feel the heat and hold it, then paraphrase and put it to the test. You receive the correction, and you design forward together. Running all four concedes nothing about whether you over-decide; the moves ask you to find out what the witness actually saw, which was never quite the thing the heat said was being alleged. The failure version is one move long: “that’s not what happened.” It feels like accuracy. It is where the data stops.
Going upstream
In-text:Chris Argyris,Overcoming Organizational Defenses (Model I/II, the three-step, skilled incompetence, and the ladder of inference). Robert Kegan and Lisa Lahey, Immunity to Change, for the developmental side. Scott Snook, Friendly Fire, for the organisational-level version and practical drift. Peter Senge, The Fifth Discipline Fieldbook, for Rick Ross’s rendering of the ladder.
Also touched: William Noonan, Discussing the Undiscussable, in the Argyris/Schön/Action Design lineage (“the problem is not me, but you” and the four thought enablers). Sidney Dekker, Drift Into Failure, on locally rational drift toward the boundary. Amy Edmondson, “Psychological Safety and Learning Behavior in Work Teams” (Administrative Science Quarterly, 1999), for the seven-item scale two of whose items the witness-round gate borrows. Mrinank Sharma and colleagues, “Towards Understanding Sycophancy in Language Models” (Anthropic, 2023), for the finding that five frontier assistants all run sycophantic and that matching a user’s view predicts which answer human raters prefer.
Go deeper: Charles Perrow, Normal Accidents, supplies the structural precondition: systems that are both interactively complex and tightly coupled turn small local adaptations into catastrophic outcomes. The U.S. Marine Corps MCDP-1 Warfighting supplies the doctrinal version in full: the local adaptation built in on purpose before it has a chance to become drift.
Watch:Roderic Yapp, Double-loop learning: a case study from the front-line (TEDxWandsworth, 17 min). Two front-line case studies, Afghanistan platoon houses and a mental-model challenge, that name the espoused/in-use gap directly. Companion: Chris Argyris, Chris Argyris Talks About Culture and Management (4 min). The framework’s primary author in his own voice, defining theory-in-use and running the canonical bypass the threat, cover up the bypassing, cover up the covering up sequence.
That gap between what you say you value and what you did under pressure sits in the leader. The same gap opens inside the change programmes you run. You sign the policy, deliver the training three times, stand up a scorecard, and still the one behaviour it all exists to produce, engineers flagging risk early, won’t hold. Every tightened deadline punishes it. Careful practice goes first. Chapter 6 asks what your organisation does to the engineer who stops the line the Tuesday the plan goes before the CEO.
Continue reading:Chapter 6 asks what your organisation does to the engineer who stops the line the Tuesday the plan goes before the CEO.
I work with engineering leaders on exactly this kind of paradigm work, the deeper the better. If it’s live for you, I’m happy to talk: schedule a 30-minute virtual coffee at hi.chrisgagne.com.
Some book links here are Amazon affiliate links; if you buy through them I may earn a small commission, at no cost to you.
“Keep everyone busy” feels like a responsible management goal. In high-variation work like software development it makes the system much slower, and the math behind that is teachable.
Chapter 3 asked whether AI is a tool or an impetus for redesign. This chapter explores whether the system around it can actually move.
🎧 Prefer to listen? This chapter is narrated in my own voice, via ElevenLabs on Spotify, about 17 minutes. Listen on Spotify →
The two poles
Pole A: utilisation. Keep resources busy. Any idle capacity is waste. Efficiency and output reports are the management metric. Manage costs first. Pole A leaders allocate every engineer to multiple projects and ask, are you fully booked? They fill an engineer’s calendar at 100% before the sprint starts.
Pole B: flow. Value throughput is the goal. Queues are the enemy: they limit our adaptiveness, sapping NPS and revenue because we can’t respond quickly to customer needs and changes in the market. Capacity margin is a feature.
Predictable work: a deterministic process, fed a steady stream of identical jobs, which queues only past 100%. High-variation work: a single queue with random arrivals and random job sizes (M/M/1), where cycle time is the work time divided by (1 − utilisation). Solid: the work itself. Hatched: time spent waiting in the queue. After Donald Reinertsen, The Principles of Product Development Flow, which draws queue size, and therefore wait time, against capacity utilisation.
Pole B leaders protect slack, accept idle time at non-constraints, and ask, “what are my queue lengths today?” as a leading indicator of lead times. These leaders don’t fill the calendar to the top. They protect 20–30% as slack, let the work pull through, and watch the queue lengths shorten.
A quick flow primer
The math is unforgiving. Donald Reinertsen, in The Principles of Product Development Flow, draws it: queue size—and therefore wait times—against capacity utilisation.
For deterministic, predictable processes—a machine fed a steady, scheduled stream of identical jobs—utilisation below 100% capacity doesn’t invoke queues. A printer that runs 30 pages a minute, fed exactly 30 a minute, never backs up; queues form only when you ask for more than 30. (Let the jobs arrive in bursts instead, and the queue returns even below 30: it is variability, not load alone, that builds the line.) This is the mental model many managers use: “meat widgets” in Craig Larman’s particularly pointed phrasing.
Stochastic, highly-variable processes—like almost all research and software development—follow an entirely different curve: a hockey stick. The average queue is governed by a doubling rule: each time you halve the excess capacity, the queue roughly doubles. The jump from 60% to 80% doubles it. From 80% to 90% doubles it again. From 90% to 95%, again. Push utilisation toward the wall and queue time, not work time, comes to dominate cycle time, and small upsets in arrivals cascade into long delays. The optimum isn’t the peak; it sits well short of it, and how far short rises with your cost of delay.
The trick is to establish stable flow first, then increase utilisation. (For another visual representation of this, watch Henrik Kniberg,The Resource Utilization Trap, about 6 minutes.)
Where Pole A is right
Genuinely fixed-throughput operations with stable demand and homogeneous work. Call centres supporting highly-regulated queries with consistent call volumes and durations. Very mature manufacturing lines on a stable specification. The conditions are narrow, but they exist.
A leader running utilisation in those conditions is reading the system correctly. With effectively zero variability—or with value-add steps so decoupled that no step waits on another—high utilisation does compound into more output.
One reaction to learning about flow is to make Pole A look stupid. This is a mistake and we can celebrate the application of Pole A in the applicable contexts. The pin factory worked 250 years ago, the 1980s call centre worked (and still does), and the high-volume assembly line works today. The leader who came up running utilisation reports came up under conditions where the reports were the right report.
The reality test is the same as in chapter 3. Is the work actually pin-factory work? For most software-dependent organisations’ portfolios in 2026, no. Some of it is. Most of it isn’t.
Where Pole B is right
Pole B better fits systems with variable demand, heterogeneous work, or genuine complexity: most knowledge-work organisations, product-development pipelines, and engineering departments. The work moves the constraint around rapidly, so there is rarely a single fixed station to keep busy (chapter 8 shows why the constraint still governs even as it moves). What you manage instead is flow.
This is where Pole A’s intuition is most exposed. The leader sees utilisation and velocity as independent variables, when the relationship between them is the curve from the primer: mathematical and unforgiving. By Reinertsen’s running tally of his conference audiences, nearly everyone measures cycle time and almost no one measures queues. Queues are the one number the curve says governs the rest.
Idle isn’t waste: a pizzeria example
I often give the example of a pizzeria. The oven fits one pizza and takes 15 minutes to bake it. The chef needs no more than five minutes to remove, cut, and box the hot pizza and prepare the next one. So 10 of every 15 minutes, the chef has nothing to do at the oven. It would be clear to any manager not to expect the chef to stay constantly busy preparing pizzas for an oven that can’t bake them at the same rate; that is wasteful. Virtually any other activity, including remaining idle, would deliver more net value. Better still, point the spare capacity at the bottleneck (perhaps the chef learns a little engineering and doubles the oven’s capacity in their spare time), and you get the most value of all. Trust the chef with the slack.
Un-baked pizzas pile up as inventory, which is to say frozen cost, not value. What’s your backlog of half-specified roadmap items, if not the same pile?
Cost of Delay is the economic argument
The sharpest move I know for a Pole-B leader in a Pole-A org is to quantify the cost of delay. Reinertsen’s E3 principle is the canonical statement: if you only quantify one thing, quantify the cost of delay. Every queue carries a cost of delay, and the cost of delay is what makes utilisation expensive. Worse, when I went looking, no off-the-shelf tool could estimate and track the cost of delay in a software context, so I built one for my own client work. Even so, most organisations still lack the means.
Every queue carries a cost of delay, and the cost of delay is what makes utilisation expensive.
The reason the argument flips intuition is that Pole A’s intuition is correct about a different variable. The cost of an engineer’s idle hour sits on the P&L as a line item. The cost of an extra week in cycle time is invisible until you put a dollar amount on it. The dollar amount, when you actually run the calculation, is usually multiples of the apparent saving from “no idle time.” Reinertsen put numbers on that asymmetry as a McKinsey consultant in 1983: a product six months late to market lost 33% of its five-year profit, while shipping on time at 50% over budget cost about 4%.
The dollarisation is what wins the boardroom: the flow argument loses with most CFOs as queueing theory and wins as a cost-of-delay number in dollars per month. One of my clients was carrying a cost of delay about an order of magnitude larger than the cost of the work itself.
Would you rather cut costs by 10%, or raise them by 10% if it brings revenue forward and grows it by more? The value reaches the market sooner, and you have longer to exploit it. The upside needn’t even be direct revenue; it can be learning that arrives sooner and compounds into revenue later. The balance is lopsided, and the lopsidedness comes from where the money sits. BCG’s 2019 analysis of 35 software companies put median R&D spend among high-growth firms at 26% of revenue. Even a heroic 50% cut to that spend, which guts the capacity that builds the upside, claws back about 13% of revenue in cost, a hard ceiling, while the upside from shipping sooner carries no such cap.
The opportunity cost compounds even further when work is completed in parallel, not serial. The utilisation reflex does more than keep individuals busy; it keeps many projects in flight at once, and that high work-in-process is its own tax. Take three equal-sized initiatives, one team, the same six months. Run them the Pole-A way, all three in flight, everyone busy, and you defer the value of all three: nothing ships until the end, so nothing earns until the end. Finish them one at a time instead and the same work, from the same team, delivers several times the value by the same date. The gain comes mostly from finishing one thing at a time; sequencing by cost of delay adds to it.
The gain comes mostly from finishing one thing at a time; sequencing by cost of delay adds to it. Equal initiatives: each is two months of the team’s work and worth the same once shipped. Delivered value counted in months, to the end of month 6. No cost counted for switching between them. Solid: the work. Hatched: waiting. Blue: delivered value. The chapter’s own example.
The AI overlay
Krivitsky, Larman, and Flemm call this the Ferrari Effect in 10X Org: picture each piece of work as a Ferrari built for speed, then jam them all on a one-lane road behind the same trucks. More Ferraris just means more traffic. AI hands every engineer a Ferrari; it doesn’t widen the road. The road is your structure, the handoffs and dependencies that decide whether work can reach a customer without waiting on someone else. Drop AI into a structure full of waiting and it amplifies the dysfunction already there. Their sequence is the whole point: first design, then AI.
When AI speeds the coding and leaves the constraint and the handoffs untouched, lead times barely move: a little faster—maybe—but nowhere near enough to earn back the token spend, because the gain hits one node in a system already saturated everywhere else.
Further, the tool the leader consults will back the approach. Ask an AI agent how to capture the productivity gain, framed as a question about individual output, and the answer usually comes back in the paradigm it was trained on: load every engineer, raise utilisation, fill the queue. The advice is fluent and confident and aimed at the pole the leader already leans toward, because as we learned in chapter 1 the LLM’s training data reverts to the mode—not the upper deciles—of human knowledge.
The longest-lived, most fashionable queue in most software organisations is the one nobody calls a queue. It’s called the roadmap. A roadmap becomes a damaging queue when stale, committed items are treated as obligations rather than revisable hypotheses. Craig Larman, in his Certified LeSS Practitioner course, names a proxy for an organisation’s adaptiveness: the percentage of items in Sprint Planning and Product Backlog Refinement that didn’t exist before the last Sprint Review. An organisation adapting to its situation produces a high number against that metric; an organisation still working a years-old roadmap produces a low one. What most software companies call a roadmap reads as a vestige of annual planning. I ran the roadmap for StubHub’s recommendations engine and read it as diligence. The one context where I owned the whole bet—my own startup—had no roadmap at all. Ryo Lu’s no-fixed-roadmap line in chapter 3 wasn’t bravado. It was a Pole-B leader naming the queue out loud, and celebrating its clearing.
The diagnostic move
Three questions for last Tuesday’s allocation meeting.
Which pole was I claiming? If asked, would I have said I’m optimising for flow or for utilisation?
Which pole would the last sprint actually show? If a colleague counted how many engineers had no slack, would the number be small (Pole B) or close to all of them (Pole A)?
Which pole does this work actually require? Heterogeneous, ambiguous, learning-heavy work calls for Pole B. Routine high-volume work with stable specification calls for Pole A. Most knowledge-work portfolios contain some of both, and the proportion is the question.
Now that you have seen the relationship between utilisation and throughput, the question is what you do with it. Not what you would say in a strategy session, but what you reward when the quarter is measured. Most leaders never chose Pole A; they have simply never priced the queue, so the resource-management spreadsheet decides for them.
If you are reading this as the CTO or VPE thinking about your CEO: the move that matters on this axis is at the metrics layer your CEO controls. Put a single page in front of them showing the cost of delay on one current initiative, in pounds or dollars, dollarised against the quarter. Reinertsen’s point is that the variable that matters most is invisible in the default management report. Make it visible in the CEO’s unit of account. In my client work, the restructuring decision has moved when the cost of not restructuring showed up in the number the CEO briefs the board on. Chapter 3’s rule holds here too: a number gets you in the door, but the win is a named decision with an owner and a date, not the number itself. The sentence your CEO can carry to the board: “We have been paying for idle-time savings with delay costs many times larger.”
The exercise: sum the queues, then price them
If you summed the total queue lengths sitting across all your backlogs right now—the tickets parked between stages, the work that entered the system but never left it—what number do you get? If you can, adjust for planned and unplanned work, and for how much your delivery rate swings iteration to iteration. Most leaders have never added it up, because almost nothing totals it for them. An estimate may be the best you can achieve.
Second: convert it into a realised cost of delay. Take one initiative sitting in that queue, put a pounds-or-dollars figure on a week of delay against it, and multiply across the wait. Expect the figure to dwarf the idle-time saving the utilisation report is protecting. When I’ve run this with clients, it produces the same silence-then-ooooh a live demo does, because the variable nobody has dollarised is the variable nobody has been managing.
In-text: the references named in the body. Donald Reinertsen, The Principles of Product Development Flow (the canonical: the three curves, the cost-of-delay E3 principle, the 100-percent-measure-cycle-time / 2-percent-measure-queues tally). The adaptiveness proxy metric comes from Craig Larman’s Certified LeSS Practitioner course. Reinertsen’s 1983 McKinsey analysis, published as “Whodunit? The search for the new-product killers” in Electronic Business, is where the six-months-late / 50-percent-over-budget asymmetry comes from. The R&D-spend benchmark is Boston Consulting Group, How Software Companies Can Get More Bang for Their R&D Buck (November 2019), whose analysis of 35 companies puts median R&D spend among high-growth software firms at 26% of revenue. The Ferrari Effect is from Alexey Krivitsky, Craig Larman, and Roland Flemm, 10X Org: Powered by Org Topologies.
Go deeper: the Theory-of-Constraints lineage this chapter rests on, developed in full in chapter 8: Eliyahu Goldratt, The Goal (don’t maximise utilisation everywhere). The software translation: Jez Humble and David Farley, Continuous Delivery (the deployment pipeline as single-piece flow), and Gene Kim, Kevin Behr, and George Spafford, The Phoenix Project (Goldratt staged in IT operations, with Brent as the constraint made personal). The cost-of-delay and dollarisation lineage: Mary and Tom Poppendieck, Lean Software Development and The Tyranny of “The Plan”; Stephen Fox and Richard Gregory, The Dollarization Discipline. The instrumentation that makes the queue math visible to non-specialists: Daniel Vacanti, Actionable Agile Metrics for Predictability (cumulative flow diagrams, cycle-time scatter plots, service-level expectations as probabilities). The doctrinal tempo extension, developed further in the leader-leader instalment: USMC, MCDP-1 Warfighting (tempo is itself a weapon; the main effort / Schwerpunkt); Robert Coram, Boyd: The Fighter Pilot Who Changed the Art of War (the OODA fast-transient theory). The further software-engineering translation: Gene Kim, The DevOps Handbook, and Gene Kim and Steven Spear, Wiring the Winning Organization (the Three Ways of DevOps); Forsgren, Humble, and Kim, Accelerate (the four delivery metrics; Ron Westrum’s pathological / bureaucratic / generative culture typology, taken up again in the batch-and-flow instalment).
The utilisation trap is a gap between the efficiency we claim to want and the busyness we actually reward. Chapter 5 finds the same gap inside the leader. Chris Argyris called it the distance between espoused theory and theory-in-use, between what you say you believe and what you visibly did when the stakes rose. Every other axis in this series lets you sort yourself into the pole that flatters you. This one doesn’t, because your team watched you decide. They have better data.
Continue reading:Chapter 5 finds the same gap inside the leader: the difference between the values you state and the ones you run on under pressure.
I work with engineering leaders on exactly this kind of paradigm work, the deeper the better. If it’s live for you, I’m happy to talk: schedule a 30-minute virtual coffee at hi.chrisgagne.com.
Some book links here are Amazon affiliate links; if you buy through them I may earn a small commission, at no cost to you.
The cheap mistake is the tokens. The expensive one is the org design you left in place.
Chapter 2 named the meta-skill of holding a paradigm as an object. This chapter is your first chance to practise: use it on your AI rollout itself.
🎧 Prefer to listen? This chapter is narrated in my own voice, via ElevenLabs on Spotify, about 28 minutes. Listen on Spotify →
The two poles
Pole A: AI as productivity layer. AI is a tool that can make existing tasks faster and cheaper. The organisation chart, role boundaries, and career ladders stay. Capable people get more done. Pole A treats AI as an adoption exercise: roll out training to start, identify the laggards, then conduct layoffs to capture the gain as a smaller team hitting the same OKRs.
Pole B: AI as impetus for org redesign. AI surfaces a question the old org chart had already stopped answering well: which work is human. Pole B hires intelligent, independent-minded, liberal-arts generalists who set their own direction, people who can define an outcome, specify the constraints, choose which work goes to AI and which stays human, and verify what comes back. Good taste now matters as much as execution skill. The specialist-scarcity logic only holds while expertise stays scarce, hard to get, and in unlimited demand for that exact skill, and Krivitsky et al. argue in 10X Org that AI dissolves exactly that. Pole B treats the generalist as the new central role.
Where Pole A is right
Pole A is right where the work is genuinely procedural, or where regulatory constraints prevent role redesign inside the planning horizon.
Pole A also fits a specific kind of work. Chapter 1 walked Adam Smith’s 1776 pin factory: identical units, effectively unlimited demand, hyper-specialisation that compounds. AI in a pin-factory context behaves like Pole A predicts: the role exists for AI to make faster and still works after AI makes it faster. The leader doing Pole-A pin-factory work and getting Pole-A pin-factory gains is doing the right thing in the right context.
The reality test for Pole A is whether the work the AI is making faster is actually pin-factory work. Most of the work running through a software-dependent organisation isn’t.
Where Pole B is right
Pole B is right wherever AI has already moved past your org chart’s design assumptions, especially where competitors are redesigning or newly forming. That describes most software-permeated (or software-vulnerable) industries right now.
Krivitsky et al. argue that nobody needs 1,000X more databases. When AI dissolves specialist scarcity, the existing org chart—not the technology budget—becomes the binding constraint. The Pole-A reflex is to consolidate—one “hyper-efficient” DBA for the whole org looks more attractive than ever—which still leaves the org chart in place, just bent the wrong way. Pole A leaders may discover two years late that their biggest cost curve was the incumbent org design itself, with token spend a rounding error against it. By then the redesign is no longer cheap, the Y Combinator startup looms larger than ever, and the window closed while managers were looking at the per-developer velocity reports. Most of what software touches now has a low barrier to entry. This context is so novel that I wonder if any prior moats will be durable unless incumbents build on them with an adaptive org redesign first.
The 10X Org operating-model material describes what the redesigned org chart looks like in practice: broader mandates, smaller teams, more outcome-orientation at the team level, less hand-off and inventory between teams. Jay Galbraith’s Star Model is underneath: strategy, structure, processes, rewards, and people have to move together. Pole A redesigns collapse to terminology and tools, leaving Galbraith’s five points largely where they were. Pole B leaders accept that the AI shift exposes all five points at once.
What the past 18 months show
10X Org made its case in the future conditional: AI-native organisations would outcompete incumbents constrained by industrial-age headcount logic. Among the leaders, that future has already arrived. Anysphere’s Cursor reportedly reached roughly $2B in annualised revenue by early 2026 with a headcount in the low hundreds. The Lean AI Native Leaderboard, on its 2025 reading, benchmarked the top ten AI-native firms at $3.48M of revenue per employee against $610K for the top ten classical SaaS companies, a 5.7x gap among the leaders. Take each figure as a timestamped illustration: the Cursor headcount, the leaderboard’s multiple, the formation rate will all move quarter on quarter. I don’t expect the direction to reverse. New formation as of early 2026 is following the same line, with roughly half of one recent Y Combinator batch building AI-native or AI-enhanced software.
The operating playbook matches what 10X Org prescribed: PM work spread across the builders, no fixed roadmap, the world changing faster than any plan can hold. Sierra’s Bret Taylor puts the incumbent failure mode in one phrase: large companies fail at AI because they are shipping their org charts. The mechanism is Conway’s 1968 law: organisations build systems that mirror their own communication structure. The case that the structure has to change first is this chapter’s, built on that law.
AI capability is jagged
Karpathy’s jagged intelligence names a paradox: the same system operating as a genius polymath in mathematical reasoning and a confused grade schooler in common-sense spatial reasoning. Tasks that feel equally hard to a human can sit on opposite sides of an invisible capability line, and you can’t know which side a task is on without trying it. The redesign question becomes which jagged edge fits which work? That question has to be asked task by task, with empirical testing, in the actual context the work runs in, and the answer changes quarterly.
Karpathy’s vibe coding names the mode; the trap is what you do with it: build by prompt, trust what comes back, forget the code exists. He frames it as floor-raising, and explicitly distinguishes it from agentic engineering, where—his words—you’re still responsible for your software just as before. The Pole-A trap is vibe-coding work that needed agentic engineering: the loose, sketch-driven build ships, and the plausible code didn’t match the actual contract.
Klarna made the same edge-misjudgement at scale, and it is one dated instance of the pattern, not the whole case: it walked back a sweeping AI-for-support move in 2025, its CEO conceding the company had leaned too hard on cost and efficiency and that quality had suffered, and it began rehiring humans for the cases where AI parity hadn’t held. The capability looked sufficient at the edge it was tested against, but the work it was actually meeting lived a few inches further in: the company had automated the cost away without automating the work.
Pole B done well holds the jaggedness explicitly. It asks, for each piece of work being redesigned: which jagged edge are we standing on, how do we test that empirically, and what is our recovery plan when the jagged edge moves?
A worked example: the inverse of NUMMI
GM tried to copy Toyota’s production system at Fremont in the 1980s through the NUMMI joint venture, sending waves of workers to Toyota City in Japan for training. The visible practices on the floor could be photographed and presumably copied in other plants. The invisible system—the supporting management functions, the supplier relationships, the relationship between the Team Member Handbook and daily work—couldn’t be seen, let alone readily transplanted into GM’s existing structure. Ernie Schaefer, the GM plant manager who later tried to replicate NUMMI elsewhere and failed, came to understand why only later, and described it on This American Life’s“NUMMI” episode in 2015:
“You know, they never prohibited us from walking through the plant, understanding, even asking questions of some of their key people. You know, I’ve often puzzled over that—why they did that. And I think they recognized we were asking all the wrong questions. We didn’t understand this bigger picture thing. All of our questions were focused on the floor, you know? The assembly plant. What’s happening on the line. That’s not the real issue. The issue is, how do you support that system with all the other functions that have to take place in the organization?”
Toyota let GM walk the plant freely because the answer GM needed wasn’t on the floor; it was in the supporting system, and GM wasn’t asking about that.
The Pole A leader is asking the GM-1985 question about AI: where on the line can we deploy this? The Pole B leader is asking the question NUMMI itself was an answer to: what does the org have to become for this to meaningfully compound?
Notice, too, how much easier Pole A is to delegate. The CEO can hire a trainer, brief the senior and middle managers, and stand up a rollout without ever changing their own week, let alone the whole org. Pole B has no such proxy. Deming put the responsibility where the systems are: management owns the systems, and therefore owns most of what goes wrong inside them. The redesign reaches the strategy and structure only the CEO owns, so it is the CEO’s own work, and far more of it. That a pole can be delegated at all is a tell that it isn’t touching the system. But in generosity to the CEO the pull toward Pole A isn’t weakness; it is the rational preference for the move you can hand off.
I find this example particularly helpful because it is a case of Pole A failing not for lack of effort or for lack of access. The GM workers were on the floor at Toyota. Toyota answered their questions. The transplant beyond NUMMI largely failed anyway, because the answers were sitting in the supporting system, not on the plant floor, and GM’s supporting system was Pole A’s supporting system. Jeffrey Liker, who has studied Japanese auto manufacturing since the 1980s, was asked on the same episode why GM’s senior management wouldn’t accept that Toyota built better cars: “I think there was pride and defensiveness. I’m proud because I’m the biggest automaker in the world. I’ve been the best. I’ve dominated the market. You can’t teach me anything, you little Japanese company.” The resistance wasn’t only passive. Dick Fuller, who ran IT at NUMMI, remembers a GM manager who toured the plant, went home, and wrote a report that amounted to “won’t work here.” “Part of that,” Fuller said, “was a threat to him. It was a threat to him to see that it was working so well.” This leads right back to Weinberg’s idea in Chapter 1: people can fold 10% into their mental category of “no problem,” whereas anything larger would be embarrassing if the consultant actually pulled it off.
Another change-agent, Jeff Weller, sent in the 1990s to convert plants, was simply thrown out of one: “I was asked in one plant to leave, because they were not interested in what I had to sell.”
The plant manager was king, and the CEO wouldn’t overrule him. The same holds now for AI inside legacy org designs: the tool is on the floor, the supporting system is the org design, and the org design keeps doing what it was built to do.
I suspect most software-dependent organisations in 2026 are still running the GM-1985 playbook: on the floor asking Schaefer’s wrong questions about AI, while the supporting system—the org chart, the career ladders, the approval gates—keeps the existing pattern intact by design.
What your last Tuesday actually shows
GM’s managers walked the NUMMI floor asking the wrong questions because they had already decided what the factory was for. The answers they got were locally correct and systemically useless. The same happens with AI procurement. You can be rigorous about what you bought and still be answering local-optima questions when the context is asking for system solutions.
I built a tool to sharpen incident learning, and it does that: it can articulate a systemic issue more precisely than most review meetings manage. What it cannot do is give a team the authority to change the conditions it has just named. The constraints that mattered were usually well out of reach of anyone in the room, or within one degree of it.
So: three questions for last Tuesday’s AI decision:
Which pole were you claiming? In the deck, the procurement memo, the rollout email, which pole did the language point toward?
Which pole would the implementation actually show? Three months out, does the team (or team-of-teams) become more adaptive, with a broader work mandate, fewer hand-offs, easier to redirect? Or just more efficient at the same work? If it is the second, Pole A delivered what Pole A promised. That describes what was chosen, and it is worth reading as data rather than as a failure.
Which pole does this situation require? This is the Cynefin question. Complex work whose consequences arrive in months calls for Pole B, the redesign. Complicated work with a stable specification makes Pole A a reasonable fit. Most senior-leadership portfolios hold some of both, so the strategic question is the proportion across the whole portfolio, one tool at a time.
The most visible gap sits between question one and question two. Leaders claim Pole B in the deck and ship Pole A in the implementation, and the team sees only the second. That gap is recoverable once you name it: you ran an efficiency play, it worked, and here is what it taught you about the redesign question underneath. The more consequential gap sits between either of those and question three. A leader whose claim and procurement both say Pole A is internally consistent, and may still be inconsistent with a context that has already moved.
The tells live in the procurement spreadsheet and the headcount plan. Pole A signs against hours-saved-per-role tools and reads efficiency dashboards, with an implicit objective of a payroll reduction that dwarfs the rise in token spend. Pole B puts its money into hiring multi-skilled generalists ahead of tool-rollout training, and makes new roles official before the AI capability is fully stable. Both are coherent. Where your three answers diverge is the next conversation to have with the CEO.
For the CTO or VPE: the decision on this axis sits at the strategy and structure layer the CEO owns. What to put in front of them is the procurement spreadsheet and the headcount plan side-by-side with the Pole A / Pole B claim in the last board deck. Put both in front of the CEO at once and ask them to read the gap. Run the same check one layer down on yourself first: do your own hiring plan and engineering ladder ship the pole your strategy deck claims? The gap lands harder upward when you have already closed it in your own house.
Before the meeting, read its history: what happened to the last person who brought this CEO unwelcome news? If the answer is bad, that pattern is the first problem, not this chapter’s. Once you’re in it, go first—disclose your own gap on the same axis, before the CEO’s—with the norm said aloud: nothing here gets used as a club, by either of us. Ask for their account before the artefact comes out. If the CEO defends the gap rather than looking at it, treat that as data rather than proof: test whether you are disagreeing about the facts, the stakes, or the frame itself. Describe what you see, flatly, then stop. As you close, end with a named decision, an owner, and a date; agreement without a date isn’t yet a win.
And if you are the CEO reading this over a colleague’s shoulder: your half is the receiving side. Ask for the story behind the deck, thank whoever went first, and hold the no-club norm yourself. The safety is real only when the most powerful person present refuses to weaponise what’s said.
The CTO’s hiring plan is where this shows up at the role level. Pole-A hiring backfills narrow specialists and adds AI tools: more frontend engineers, more backend engineers, each productivity-boosted. Pole-B hiring inverts it: a smaller team of generalists, each able to define outcomes and orchestrate AI rather than execute against a specification. By the time the redesign window closes, the Pole-A CTO has built a bench of narrow specialists AI is dissolving, and now has to retrain them, lay them off, or watch the AI-native competitor recruit them.
The question to bring is which roles in this company still depend on the narrow-specialist assumption AI is dissolving? For scale, as of early 2026 the top AI-native firms are running several multiples ahead of their classical peers on revenue per employee. This is a CEO decision about what the org is, not a technology decision about what the tools do. The sentence your CEO can carry to the board: “We have been investing in AI tools and not in the org design that makes them pay.”
An exercise: the shadow org chart
Lay the shadow chart beside the real one. The gap between them is your own estimate of how far the redesign has to travel, and you built it without a meeting.
Start with a move you can make tomorrow, alone, without anyone’s permission: take one engineering team or team-of-teams and redraw it on a blank page as if AI capability had been assumed from the day the team was formed. Ignore the current titles, skill sets, and reporting lines. Sort the team’s actual work into three bands: AI-autonomous, where a human reviews outputs but doesn’t generate them; AI-supervised, where a human directs the AI and verifies what comes back in real time; and fully human. Then draw the team that work implies: how many people, in what roles, reporting to whom. Lay the shadow chart beside the real one. The gap between them is your own estimate of how far the redesign has to travel, and you built it without a meeting.
Then ask the second question. For each role that survives, what does the human job on the other side actually look like? Map it: title, career ladder above and below, hiring filter, on-ramp. Expect one of two results. Either the new role doesn’t exist yet in the org chart, and the redesign work begins in earnest. Or it exists, but sits inside a career ladder built for the old role, with progression criteria that still reward the old behaviour. Roles are downstream of strategy, structure, processes, and rewards; the map makes the downstream relationship visible. Be mindful not to lean on existing skill-based roles; consider what skilful generalists with good taste can do.
Then make it shared. Run the map a second time with the CEO against the org chart, without the team present, and name who decides what changes on the back of it and on what timeline. If the answer is the role redesigns, bring it to the team and redo the map with them: the leadership-only version is a hypothesis about work-as-imagined, and chapter 14 explains why it will be wrong about work-as-done. If the answer is the work has dissolved, that is a different conversation that doesn’t belong in this exercise. The team is owed the truth, not an exercise that arrives looking like an ambush. The solo shadow chart is yours to run tomorrow; the shared redesign is a CEO decision about strategy and structure, and the CTO carries the role-bifurcation half.
The cheap mistake is the tokens. The expensive one is the org design you left in place.
Sources for the AI buyer who wants to go upstream
In-text: Krivitsky et al., 10X Org: the structural case behind this chapter, including the specialist-scarcity claim and the nobody needs 1,000X more databases argument. The paradigm-and-identity extension is mine. The Lean AI Native Leaderboard on the 5.7x revenue-per-employee gap. Ernie Schaefer’s account of why the NUMMI transplant failed, and Jeffrey Liker on the pride and defensiveness he saw in GM’s senior management, both from This American Life’s “NUMMI” episode (2015). Andrej Karpathy’s 2025 LLM Year in Review for jagged intelligence, vibe coding, and agentic engineering, plus his Software 3.0 framing from his Sequoia AI Ascent interview. Bret Taylor on companies shipping their org charts.
Also touched: Jay Galbraith, Designing Organizations, on the Star Model: strategy, structure, processes, rewards, and people moving as one.
Go deeper:Watch. Craig Larman, Craig Larman’s Take on 10X ORG: Why He Thinks This Book Matters (5 min), names the dichotomy directly: Pole A as the cost-curve trap, where AI is cheap, always-on labour and a human who offers only a 10% improvement won’t compete; Pole B as the 10x redesign claim, with 10xOrg, Org Topologies, and Large-Scale Scrum (LeSS). Read. Org Topologies (start with their primer) on the operating-model maps that go with the redesign. On the macroeconomic frame, Daron Acemoglu and Simon Johnson, Power and Progress, and Acemoglu’s The Simple Macroeconomics of AI (2024 paper), whose “so-so automation” names the Klarna pattern above. Ethan Mollick’s Co-Intelligence for the practitioner-side empirical material and the jagged frontier framing folded into Karpathy above. Conway’s 1968 law, that organisations design systems mirroring their own communication structure, behind Taylor’s phrase. Cursor’s Head of Design Ryo Lu on how the work actually runs, in his November 2025 interview with Peter Yang, as reported by DevClass: a lot of the PM jobs are spread across the builders in the team, and no fixed roadmap because the world is changing faster and faster, there’s new models dropping every day. Further AI-native revenue datapoints: Sierra at $100M ARR in 21 months on outcome-priced agents, ElevenLabs past $500M ARR in early 2026 with about 530 people, and the W26 batch analysis counting half the companies as AI-native or AI-enhanced. Klarna’s walk-back: it had replaced the equivalent of some 700 support agents with an AI assistant before reversing in 2025.
If the expensive mistake is the org design you left in place, that same design is deciding how fast work moves through it. “Keep everyone busy” feels like responsible management. In high-variation work like software it slows the whole system down hard, and the maths that proves it takes minutes to learn. Chapter 4 shows you that curve, and where your own org chart already sits on it.
Continue reading:Chapter 4 shows you that curve, and where your own org chart already sits on it.
I work with engineering leaders on exactly this kind of paradigm work, the deeper the better. If it’s live for you, I’m happy to talk: schedule a 30-minute virtual coffee at hi.chrisgagne.com.
Some book links here are Amazon affiliate links; if you buy through them I may earn a small commission, at no cost to you.
Think of a physiotherapist. They don’t fix your back. They put a mirror in front of you, show you the exact posture you’ve been holding for fifteen years, and say: “notice you’re doing that? that’s why you’re in pain.” The pattern was invisible from the inside. Once it’s visible, the work is small corrections, practised many times, especially under load. This piece is that kind of mirror.
NB. The dichotomies below are mostly not original to me; the curation and the synthesis are. I name teachers in each entry and link to their work so you can go upstream when something resonates. If this mirror surfaces something worth working, the people whose names appear in the Watch and Read lines are where I’d send you next. Or me—this is the work I do.
When senior leaders see the operating paradigms they have been running on, the calls they make under pressure stop surprising them. With that, the response that better fits the new situation becomes available. The work I do with leaders suggests most are running paradigms they cannot see, built for problems the firm no longer has, in territory the paradigms weren’t built for. AI is the most visible driver of the changing context, but not the only one. The piece below offers a mirror for seeing what’s running underneath: 28 observable decision-level dichotomies, written as prompts to compare what you say you believe to what you actually do.
This is the canonical long-form treatment of the work; a shorter, axis-by-axis article series will follow for readers who want to take it in smaller doses.
Who this is for, who it isn’t. This is for senior leaders of software-dependent firms whose situation-mix has shifted toward complex, unpredictable work: long-term survival under uncertainty, AI-era operating-model threats, leading technology-driven market disruption. You might suspect the mindset that got you where you wanted to be yesterday isn’t fitting what your new context needs today, or you might find yourself in the midst of a crisis. It isn’t a leadership-style typology, a developmental ladder, or a self-administered test. You are not broken and do not need fixing; you are simply running on paradigms that fit a context you have since left. If your work is genuinely and predominantly Complicated/Ordered (regulated compliance, financial close, structural engineering, mature manufacturing on a stable specification), the classical management paradigms are possibly sufficient. That said, almost no senior role in a software-dependent firm is purely Complicated/Ordered anymore: take thirty minutes anyway, and notice how much the circumstances have shifted underneath you.
The diagnostic question. Most knowledge work has been Complex/Unordered for decades: the territory where the answer isn’t knowable in advance, only discoverable by probing (Dave Snowden’s Cynefin, defined more fully in “The frame” below). What changed is the cost of mismatch. The Complicated/Ordered habits—the classical management toolkit, built for work where experts can analyse the problem to a known good answer—could carry a slow-moving firm through a Complex/Unordered situation as long as the consequences of being slightly wrong arrived slowly too. With AI accelerating the consequences and reshaping the substrate, the gap between the leader’s habitual response and what the situation actually requires now compounds in months rather than years. So the diagnostic question is: my role has been Complex/Unordered all along; is my habitual response still fitting now that the cost of mismatch is, for a growing number of firms, terminal?
On Chegg’s May 2023 earnings call, CEO Dan Rosensweig named ChatGPT as the threat to new-customer growth; they launched CheggMate, their OpenAI-built tutoring product, the same month. Three years later the market cap is down roughly 99% from its 2021 peak, headcount has fallen from ~3,200 to under 600, and the Q1 2026 homework business is down 57% year-on-year. They posted their first profit in two years by cutting costs faster than revenue fell. Awareness was not the limiting factor; the org could not redesign around the substrate shift in time.
The three most important paradigm dichotomies for most readers in 2026. If you read nothing else, read these:
Entry 1, AI as superficial productivity layer vs AI as substrate for org redesign. The dichotomy that decides whether the firm’s AI investment compounds or runs out as a local productivity uplift that doesn’t translate to revenue.
Entry 6, Utilisation vs flow. The dichotomy that decides whether AI speeds your system up or just relocates the queue.
Entry 11, Espoused theory vs theory-in-use. The dichotomy that decides whether your stated principles survive contact with your last three decisions under pressure.
The six paradigm axes and the two poles. Each axis asks where your operating reflex sits. Pole A is the move that often feels like leadership; Pole B is what the Complex/Unordered contexts reward most. Both poles have bounded applicability, neither is correct for all contexts. The question is whether your habitual response fits the situations you actually meet.
Change: how this leader produces change.
1. AI bolted onto the existing organisational design, or the organisational design entirely reconsidered around AI.
2. Announce the change programme and run training, or redesign what surrounds the behaviour so the new move is the easier move.
3. Write the strategy and roll it out, or design the conversation that produces the strategy.
4. Ship a directive, or ship a question and an open week.
5. Send people on courses, or block thirty minutes a day for coaching at the work.
Optimisation: how this leader reads the unit of value and the shape of waste.
6. Keep everyone fully utilised, or keep queues and lead times short.
7. Penalise plan variance, or plan a hypothesis and an experiment.
8. Set efficiency targets per function, or find the one constraint and subordinate everything to it.
9. Celebrate the launch, or celebrate the customer succeeding three months later.
10. Long roadmaps and/or Planning Intervals, or continuous flow, WIP limits, and pull.
Knowledge: what this leader treats as legitimate knowledge.
11. Judge yourself by your intentions, or judge yourself by your last three decisions under pressure.
12. Project certainty, or voice a hypothesis and the test that would falsify you.
13. Run the checklist, or trust the experienced operator.
14. Standardise across business units, or protect what the local team already knows how to do.
15. Convene more analysis when stuck, or pause and check what the experienced operator already senses.
Authority: how this leader reads power, decision rights, and their own role.
16. Fill coordination channels with approvals, or fill them with “I intend to…”
17. Decide where the formal power is, or decide where the tacit knowledge actually lives.
18. Get the work out of people, or get out of people’s way.
19. Be the hero, or build up your successors.
20. Cascade the OKRs down the org, or redesign decision rights so the org can move without you.
21. Open with the answer, or open with a question you don’t already know the answer to.
22. Be the deepest expert in the room, or convene the people who know.
Causation: how this leader reads the relationship between events and their causes.
23. Find the deviator who failed, or close the gap between the work as imagined and the work as done.
24. Count incidents, or study what works well (not just what broke).
25. Tell one clean story, or hold several stories at once.
Self: how this leader reads their own identity, capacity, and developmental edge.
26. Hold the role as who you are, or hold it as what you’re currently doing.
27. Be surprised by your own pattern when it surfaces, or name the anxiety driving it.
28. Defend the decision after it fails, or refine your mental models after they fail.
15 minutes today. Look back at the six axes above. For each axis, notice which pole you sit closer to. Where you sit on the first pole, ask: where did this assumption come from, and is it still fitting the work you actually do now? That alone will be enough to start the noticing the rest of the article is built around.
If the 15-minute pass surfaces something worth more than a glance, continue reading and work the two-week protocol in How to use this mirror at the end.
Why a mirror, and why now
Jennifer Garvey-Berger puts it plainly: “We get so used to our patterns that we can forget what the situation calls for and just rely on our own habits.” That is the diagnostic problem this mirror addresses.
A senior leader who has read this far is still left with a practical question. Which paradigm am I actually running? The intellectual answer is almost always the flattering one. The honest answer lives in last Tuesday’s meeting, in the calls I made under pressure when nobody was watching me make them.
This mirror is ultimately meant to be used in conversation: with a coach, a trusted peer, a few witnesses willing to compare what you say to what they actually saw you do.
It shows you what habit has hidden from you. The change is the small correction—repeated under load—in the conversations that follow. One mirror is rarely enough; like a physiotherapist’s, the role is distributed across several people who will keep showing you the same pattern from slightly different angles until you can feel it yourself. You’ll need at least a consultant who understands the paradigms and a developmental coach who can help you through the change; these are rarely the same person.
The frame: the situation-mix in your role has shifted
Snowden’s Cynefin distinguishes five kinds of situation:
Clear. Best practice, fixed constraints, self-evident cause and effect.
Complicated. Good practice via expert analysis, knowable causal relationships, governing constraints.
Complex. Exaptive practice via probing, enabling constraints, emergent and dispositional behaviour (which is interestingly something an LLM also embodies).
Chaotic. Novel practice, no effective constraints.
Confused. Knowing which domain you are in is itself the question.
Cynefin is a sense-making framework for assessing situations, not for typing leaders. The same leader routinely meets situations in all of those domains and is asked to respond to each.
The diagnostic question this piece works with is therefore not “which domain do you operate in.” It is: has the situation-mix in your role shifted toward Complex/Unordered work, and does your habitual sense-making mode fit the situations you now meet?
For most senior leaders in software-permeated firms, the answer to the first half is plainly yes. Strategy under genuine uncertainty, leadership transitions, cross-functional change, AI-era organisation redesign, talent dynamics, market position under technological disruption: these are all Complex/Unordered. This has been true for most knowledge workers for decades. The classical management toolkit was built for Complicated/Ordered work—regulated compliance, financial close, structural engineering, mature manufacturing—and remains right in that domain. The trouble starts when a Complicated/Ordered habit meets a Complex/Unordered situation and treats it as if the answer were knowable.
Ronald Heifetz puts the same diagnostic in a different vocabulary:
Technical problems have known solutions accessible through current authority and expertise: forecast, plan, hold the plan as the basis for accountability; when results slip, take more direct control or buy the answer.
Adaptive problems require changes in values, beliefs, roles, or relationships, and no single action controls the outcome: run safe-to-fail experiments, hold the group in productive disequilibrium, let pattern emerge.
Heifetz’s adaptive-vs-technical and Snowden’s Complex/Unordered-vs-Complicated/Ordered are similar diagnostics at different scales; every entry below is a slice of the larger question: what does this leader do when the response that used to work no longer fits the situation?
I learned this exercise from Bas Vodde. Draw a horizontal axis. Theory X to the left (workers cannot be trusted, require coercion and control), Theory Y to the right (work can be as natural as play, people seek responsibility and meaning). Sticky-notes above the line are people’s observations of their colleagues’ attitudes and behaviours. Sticky-notes below the line are observations of the organisational context they’re operating inside. The gap between the two rows of sticky-notes, when it appears, is the data:
A redacted artefact from a Theory X / Theory Y exercise I facilitated. The top row clusters Theory Y; the bottom row clusters Theory X.
Pole A, Theory X. Design control structures around the worst-case employee; oversight as the default; trust as a reward earned through visible compliance.
Pole B, Theory Y. Design control structures around the average employee; trust the system to handle the exceptions; remove obstacles and let people do good work.
Where Pole A is right. Genuinely high-stakes regulated work where the cost of a single bad actor is catastrophic: nuclear plant control rooms, aviation cockpit checklists, surgical theatre protocols, anti-money-laundering compliance. The structural Theory X is not a failure of trust; it is the right engineering response to a Clear- or Complicated-domain risk surface in a low-variance context.
Where Pole B is right. Most creative work, most knowledge work, genuinely Complex cross-functional decisions, almost any role where intrinsic motivation is doing the real work.
What it looks like in decisions. Pole A leaders write detailed approval workflows for spending below $5k. Pole B leaders set a budget and let the team allocate.
Source. Douglas McGregor.
The larger mirror
Here are all 28 paradigm dichotomies, sorted into six roughly-organising axes.
Change: how does this leader read and produce change?
1. AI as productivity layer vs AI as substrate for org redesign.
Pole A. AI is a tool that makes existing roles faster, cheaper, or more leveraged. The organisation chart, role boundaries, and career ladders stay; capable people get more done. Pole A treats AI as adoption.
Pole B. AI is a substrate that re-shapes which work is human and which is not. The valuable human role broadens rather than narrows: the builder, or skilful generalist, who can define an outcome, specify the constraints, choose which work goes to AI and which stays human, and verify what comes back. The narrow specialist’s economic logic—expertise is scarce and demand for it is unlimited—is the assumption AI dissolves. Pole B treats AI as redesign and treats the generalist as the new load-bearing role.
Pole A is right where the work is genuinely procedural, the regulatory or domain constraints prevent role redesign within the planning horizon, or the firm is buying time to learn what to redesign. Pole B is right wherever the planning horizon is shorter than the rate of capability change in the AI substrate, or where competitors are already redesigning.
In decisions: Pole A leaders sponsor AI-rollout programmes inside existing job families and ask “what’s the productivity uplift?”; Pole B leaders ask “which work is now human, which is now AI-supervised, which is now AI-autonomous, and what does the role—if any—on the other side actually look like?” They accept that the answer reshapes the entire organisation’s strategy, process, people, structure, and rewards. The traps on both sides are different. Pole A leaders may well discover two years late that their biggest cost curve was the incumbent org design itself, not tokens. Pole B leaders get the org design right (adaptive topology, broadened mandates, smaller teams of builders) but might bet specific designs on AI capabilities that aren’t yet stable, and have to walk those back.
The builders carrying the substrate shift in the firms I see doing it well tend to share a profile: high pattern-recognition across domains, liberal-arts-educated as much as they are technically trained, comfortable holding business, technical, and human-systems vocabulary in the same sentence. They are noticeably less patient with the status quo than average. The org-design literature names the role (multi-learning, M-shaped, expert generalist) but undersells how specific the profile is. Hiring for it is harder than hiring for the role it replaces.
Watch. Craig Larman, Craig Larman’s Take on 10X ORG: Why He Thinks This Book Matters (5 min). Names the dichotomy directly: Pole A as the cost-curve trap (“AIs cheap as dirt, 24/7, no vacation, humans won’t compete by suggesting ‘hire me, I’ll make you 10% better’”), Pole B as 10x redesign rather than 10% improvement, with 10xOrg, Org Topologies, and LeSS. Companion: the Reinertsen video from entry 6 for the flow-economics half.
Read. Org Topologies (start with their primer) and 10xOrg on org-design at AI-substrate scale; Jay Galbraith, Designing Organizations (Star Model) on the model that underpins much of Org Topologies and 10xOrg; and Reinertsen, The Principles of Product Development Flow, on what happens to flow and cycle time when the speed of individual process steps changes faster than the surrounding system.
2. Pushing people vs changing the conditions.
Pole A. Produce change by force of will: announce the programme, run the training, hold people accountable for the new behaviour.
Pole B. Produce change by shifting what surrounds the behaviour: the path (incentives, friction, information flow) and/or the operating system (the meaning-making that produces the behaviour in the first place). Pole B has two flavours: environmental design (change the situation, behaviour follows) and developmental work (change the mindset, behaviour becomes durable). Change the organisation design and the behaviours will also change.
Pole A is right in time-critical regulatory or competitive shifts where the new behaviour is specifiable and adoption speed is the binding constraint. Pole B is right whenever the change is adaptive rather than technical, or whenever Pole A compliance reliably decays the moment pressure relaxes.
In decisions: Pole A leaders announce programmes and run training; Pole B leaders redesign the environment so the new behaviour is the easier behaviour, and accept that the deeper version of the work is developmental, takes much longer than scheduled, and treats early relapse as “not yet” (in Heath & Heath’s frame from Switch, drawing on Carol Dweck) rather than as failure. Training is still valuable—especially when it transmits difficult-to-document embodied wisdom—just not sufficient.
Pole A. Get the right answer; convince others; the decision itself is the work.
Pole B. Get the right way of deciding; the quality of the process determines the quality and durability of the decision.
Pole A is right when the answer is genuinely knowable and the legitimacy of the decision is not at stake: the technical call on a deployment rollback, the regulatory interpretation, the budget reallocation inside an already-agreed envelope. Pole B is right when the decision will be implemented by people who weren’t in the room and whose ownership of the decision is what makes it stick, which is most systemic decisions.
In decisions: Pole A leaders write the strategy and roll it out; Pole B leaders design the conversation that produces the strategy.
Watch. Roger Schwarz, Smart Leaders, Smarter Teams (4 min): the leader as architect of how the team thinks together, not just decider of what gets decided.
Pole A. Authority issues the requirement; people comply.
Pole B. Authority articulates purpose; people enrol or commit; compliance is the floor, not the goal.
Pole A is right when speed is binding and the cost of partial enrolment is acute. Pole B is right when the work depends on enrolment for any durability.
In decisions: Pole A leaders ship a directive; Pole B leaders ship a question and an open week.
Watch. Peter Senge, On Shared Vision (2 min): on the compliance-to-commitment ladder and why a vision has to be talked about with people, not told to them. Companion: Frederic Laloux, 1.2 What truly drives you? (Thought for top leaders) (10 min) on speaking from the why rather than mandating the what. The University Hospital CEO story shows the exact moment a leader stops mandating concepts and resistance vanishes.
Pole A. Episodic, off-site, formal; the classroom is the venue.
Pole B. Daily, embedded in work, mentor-mentee at the gemba; the workplace is the venue.
Pole A is right when the content is genuinely transferable from classroom to context and the transfer overhead is low: compliance certifications, language fluency, formal credentialing requirements, software platform onboarding for stable tools. Pole B is right when the content lives in the situated work and the classroom abstraction loses most of what matters: coaching skills, judgment under uncertainty, leadership behaviour under threat, and almost any capability the leader cares about in their senior people.
In decisions: Pole A leaders send people to courses; Pole B leaders block thirty minutes a day for coaching cycles at the work.
The exception worth naming: well-designed games, simulations, and case-based kata are training in venue but Pole B in substance. They transmit the tacit pattern-library—the mētis—that “sage on the stage” classroom abstraction usually can’t carry. Donald Schön’s practicum, Gary Klein’s recognition-primed decision simulation work, and Ikujiro Nonaka’s socialisation mode in SECI all name this third space. The diagnostic isn’t the venue; it’s whether the content is codified or situated.
Optimisation: what is the unit of value, and how does this leader read waste?
6. Utilisation vs flow.
Pole A. Keep resources busy; idle capacity is waste; efficiency reports are the management metric; manage cost first.
Pole B. Throughput is the goal; queues are the enemy; capacity margin is a feature, not a bug. As utilisation climbs toward 100%, queue time dominates cycle time and small upsets in arrivals cascade into long delays.
Pole A is right in genuinely fixed-throughput operations with stable demand and homogeneous work: call centres on routine queries, very mature manufacturing lines. Pole B is right in any system with variable demand, heterogeneous work, or genuinely }Complex/Unordered prioritisation.
In decisions: Pole A leaders allocate every engineer to a project and ask “are you fully booked?”; Pole B leaders protect slack, accept idle time at non-constraints, and ask “what are my queue lengths” as a leading indicator of lead times.
The three curves say the same thing: the optimum is not the peak. Utilisation as the metric runs you past it. Queue size is the textbook single-server curve, ρ/(1−ρ); the other two charts show shape only. After Don Reinertsen, FlowCon 2014.
Pole A. Strategy as document, vision as destination, plan as commitment; decide early, execute against the commitment; variance is failure.
Pole B. Strategy as verb; set direction, find adjacent possibles; treat each plan as a hypothesis with an expected outcome; check; act on the difference; “cross the river by feeling the stones.”
Pole A is right when the planning horizon is shorter than the rate of change, the cost of changing the plan is high, and accountability requires commitment. Pole B is right when none of those is true.
In decisions: Pole A leaders penalise plan variance and demand a guarantee; Pole B leaders demand a hypothesis and an experiment, update the plan, and treat the update as evidence the strategy is working.
Watch. Mary Poppendieck, The Tyranny of “The Plan” (60 min, transcript also available on LinkedIn; this is canonical and well worth the time).
Pole A. Improve each department on its own metrics; the whole will improve as a sum.
Pole B. Identify the one constraint that limits the whole, exploit it, subordinate everything else to it. “A system of local optimums is not an optimum system; it is a very inefficient system.”
Pole A is right only when the system genuinely decomposes: independent product lines, separate geographies, decoupled value streams with no shared constrained resource. But most systems leaders think decompose don’t. The resources, the people, the queue capacity, the management attention all run through shared bottlenecks the org chart hides. In manufacturing—a Complicated/Ordered domain—Goldratt demonstrates that local efficiency improvements on non-bottleneck resources reduced plant throughput, by pushing more work onto the bottleneck and consuming time that couldn’t be recovered. W. Edwards Deming’s funnel experiment proves the same point for stable production processes: tampering with common-cause variation makes the variance worse, not better. The claim isn’t restricted to Complex/Unordered work; it holds anywhere events depend on prior events and demand has variability, which is nearly every system a senior leader actually meets.
In decisions: Pole A leaders set utilisation targets per function; Pole B leaders ask “which one resource, today, is governing how fast the whole system can move?” and protect it.
Pole A. The unit of management is units shipped, features released, work completed.
Pole B. The unit of management is customer progress, market position, mission impact.
Pole A is right when output and outcome have been demonstrated to track each other tightly in the relevant domain: a sales organisation on a mature product where deals booked do convert to revenue, a fulfilment operation where units shipped equals customer demand met, a regulatory pipeline where filings submitted equals approvals secured. Pole B is right whenever the firm is producing more output and getting worse outcomes, a common pattern in feature-factory product organisations and in any system where the output metric has become the goal rather than the proxy.
In decisions: Pole A leaders celebrate the launch; Pole B leaders celebrate the customer using the thing successfully three months later.
Pole A. Aggregate work into larger units to amortise setup costs; schedule and dispatch work to resources based on forecast; central plan governs.
Pole B. Reduce batch size; resources pull work when they have capacity; demand signal at the customer end governs upstream.
Pole A is right when setup costs are genuinely fixed and high, forecast accuracy is high, and the cost of inventory is low (classical EOQ logic in mature physical operations). Pole B is right in any context where setup costs can themselves be reduced and the learning rate matters, or where forecast accuracy is low.
In decisions: Pole A leaders schedule long roadmaps and allocate capacity by Planning Interval; Pole B leaders ship continuously, set WIP limits, and let teams pull from a ruthlessly prioritised and pruned backlog.
Knowledge: what does this leader treat as legitimate knowledge?
11. Espoused theory vs theory-in-use (Model I vs Model II).
Pole A. The leader’s stated principles, values, and intentions are taken as the relevant data; when results miss, change the action while leaving the underlying assumptions intact. Under threat or embarrassment, Model I governs: win not lose, be rational, avoid upset, support others by telling them what they want to hear; the four social virtues that quietly underwrite organisational defensive routines. “The problem is not me, but you.”
Pole B. The leader’s actual behaviour, especially under threat or embarrassment, is taken as the data. The gap between espoused and in-use is the diagnosis. Under threat, Model II governs: grant legitimacy to others’ views, assume partiality of one’s own, attribute positive intent, acknowledge impact and contribution. “Aggressive and vulnerable”: strong advocacy paired with genuine inquiry into one’s own contribution.
Pole A is right for external communication: stated principles need to be stated clearly, and an organisation needs them on the record. Pole B is right for the leader’s own self-assessment, and for any conversation where defensive reasoning is already running and the next decision will repeat the last one unless the gap is named. Most leaders, by their own report later, find Pole A was running underneath their Pole B account in at least some of the territory.
In decisions: Pole A leaders judge themselves by their intentions, run after-action reviews on the what, and end disagreements with a winner; Pole B leaders judge themselves by their last three observable interactions, run after-action reviews on the what we believed coming in, and end disagreements by asking which observable example would change the other person’s view, and whether the same example would change theirs. This dichotomy is the load-bearing wall of this mirror: every other entry can be self-reported into the flattering pole, and only the gap between espoused and observed produces signal the leader didn’t already have.
Watch. Roderic Yapp, Double-loop learning: a case study from the front-line (TEDxWandsworth, 17 min): two front-line case studies, Afghanistan platoon houses and then a mental-model challenge, that name the espoused/in-use gap directly. Companion: Chris Argyris, Chris Argyris Talks About Culture and Management (4 min). The framework’s primary author in his own voice, defining theory-in-use and running the canonical “bypass the threat, cover up the bypassing, cover up the covering up” sequence.
Pole A. Strong, confident, resolute positions communicate competence; doubt is weakness.
Pole B. Lack of certainty is strength; certainty is arrogance. Think out loud; voice context and hunches; invite challenge; make 90% confidence interval predictions and audit calibration after.
Pole A is right in genuine emergencies—Chaos in Cynefin terms—where projected confidence is the team’s anchor and the cost of paralysis exceeds the cost of being wrong. Pole B is right almost everywhere else, especially in low-validity decision domains where confidence tracks coherence of story rather than quality of evidence.
In decisions: Pole A leaders make point predictions and defend them; Pole B leaders advance positions with the test that would falsify them.
Pole A. The right answer comes from listing the options, comparing them against criteria, and choosing the best. Deliberate analysis is what makes a decision defensible.
Pole B. The right answer comes from recognising the situation as one the experienced operator has seen before, generating one workable course of action, and mentally simulating it forward. Pattern-recognition is what makes a decision fast and, in the right conditions, reliable.
Pole A is right when the environment is irregular, feedback is delayed or absent, the stakes are high enough that the decision will be publicly justified, or the operator is not yet experienced enough for Pole B to be trustworthy. Pole B is right when the environment is regular enough that experience has built trustworthy patterns and feedback is rapid enough to keep those patterns calibrated: emergency response, surgical theatre, deployment rollback, anything an expert has done a thousand times. Klein and Kahneman, who spent years disagreeing about whether expert intuition was real, eventually agreed: it is real where the environment is regular and feedback is rapid, and it is unreliable everywhere else. That joint conclusion is the diagnostic for which pole fits.
In decisions: Pole A leaders introduce a checklist or formal review where stakes are high and feedback is poor; Pole B leaders trust the experienced operator’s first reasonable option and notice when the situation has slipped outside the territory the operator’s experience covers.
Pole A. Knowledge that matters is universal, codified, decomposable, teachable as formal discipline; standards, documentation, and knowledge-management systems are the organisation’s memory.
Pole B. Most operating knowledge is local, implicit, and held by practitioners: James C. Scott’s mētis, Nonaka’s tacit knowledge, the river pilot’s feel for the one harbour. Mentor-mentee transmission and shared experience are the real infrastructure.
Pole A is right where the work is genuinely procedural, the knowledge is genuinely codifiable, and personnel turnover is high enough to need it. Pole B is right wherever a work-to-rule strike would halt the operation, which is almost everywhere outside fully automated processes.
In decisions: Pole A leaders standardise practices across business units and invest in documentation; Pole B leaders ask the local team what they already do that the standard fails to capture, protect it, and invest in senior staff coaching junior staff at the work.
Pole A. Information is gathered, analysed, decided upon; the cognitive faculty is the relevant tool.
Pole B. The body knows before language does; felt sense, presence, embodied attunement are signal, not noise.
Pole A is right when the situation is tractable enough that analysis converges and the body’s signal is dominated by cognitive bias. Pole B is right when the situation is intractable, analysis has converged on a position the body finds suspect, and the suspicion is data.
In decisions: Pole A leaders convene more analysis when stuck; Pole B leaders pause, attend to body, name the discomfort, and let the next move emerge.
Watch. Eugene Gendlin, Focusing (Nada Lou archive, 12 min). Gendlin himself on felt sense: “the body has its own take on what’s going on… more subtle, more intricate… you walk by the familiar feelings… one more step.”
Authority: how does this leader read power, decision rights, and the leader’s role?
16. Leader-follower vs leader-leader (and the orders that go with each).
Pole A. Authority lives with the leader; subordinates seek permission and report status; orders are detailed, control is by supervision, deviation is failure.
Pole B. Authority lives where the information lives; subordinates state intent and act, the leader replies “Very well”; orders are intent, purpose, constraints, and antigoals, with subordinates choosing the right action as conditions change.
Pole A is right in genuine emergencies where time-to-decision is the binding constraint and command structure is what is being practised, or where the right action is precisely specifiable and the cost of unauthorised initiative is high. Pole B is right almost everywhere else, including most of what looks like an emergency, and especially where the right action depends on local conditions.
In decisions: Pole A leaders’ calendars fill with approval meetings and they write detailed task lists; Pole B leaders’ calendars fill with conversations that begin “I intend to…” and they write a paragraph of purpose and a list of what must not happen.
Pole A. Information flows up, decisions flow down; the org chart is the operating system.
Pole B. Information flows where it is needed; decisions are made where the information is; the org chart is one of several views.
Pole A is right where regulatory accountability or genuine command structure requires it: defence, regulated finance, certain public-sector contexts. Pole B is right almost anywhere else, including in regulated industries where the structure has been kept for inertial reasons.
In decisions: Pole A leaders re-org by redrawing reporting lines; Pole B leaders re-org by changing meeting cadence, information flow, and decision rights. Steve Jobs named the Pole B move plainly at D8 2010: Apple is “organised like a startup… run by ideas, not hierarchy.”
18. Theory X vs Theory Y. See the worked example above. Pole A is right in two distinct contexts: genuinely high-stakes regulated work with catastrophic-bad-actor risk (pharmaceutical batch release, nuclear plant control rooms, anti-money-laundering compliance, aviation cockpit procedure), and routine low-skill, low-discretion work where the task itself doesn’t generate intrinsic motivation and the structure is doing the load (high-turnover transactional roles, certain warehouse or contact-centre tasks where the work is the work and capability development sits outside it). Pole B is right in most knowledge work, most cross-functional work, and most situations where intrinsic motivation is doing more of the load than the control structure can see.
Pole A. The leader’s authority and competence drive the team; subordinates execute; the leader takes personal credit for success and personal responsibility for failure; the organisation orbits the leader’s competence.
Pole B. The leader retains final authority but creates highly participative teams; builds capacity that outlives them; the test of leadership is what happens after they leave.
Pole A is right in genuine crisis where personal visibility holds the team together, and is over by the time the crisis ends. Pole B is right almost everywhere else.
In decisions: Pole A leaders close debates by deciding, accept extension after extension because the organisation “needs them”; Pole B leaders close debates by surfacing the disagreement and inviting the team to decide, and measure success by their successor’s competence.
Pole A. Organisation as machine; meritocracy, MBO, KPIs, R&D, balanced scorecards; effectiveness as the yardstick.
Pole B. Organisation as living system; self-management, wholeness, evolutionary purpose; sense-and-respond rather than predict-and-control.
Pole A is right in tractable, stable industries with well-understood competitive dynamics: regulated utilities, mature consumer staples, infrastructure with predictable demand. Pole B is right in industries undergoing genuine technology-driven disruption where the planning horizons no longer track the disruption rate, which is most software-permeated industries right now.
In decisions: Pole A leaders run quarterly cascades of OKRs; Pole B leaders run quarterly listening practices to “what the organisation wants to become.”
Pole A. The leader’s job is to know the answer and communicate it clearly. Effective leadership is articulate direction; effective change is handing the team the answer, the plan, the playbook.
Pole B. The leader’s job is to surface the information needed for good decisions. Ask questions to which you do not know the answer; let the team find the answer; develop capability, not just compliance.
Pole A is right when the leader genuinely knows the answer, the team genuinely does not, and either the cost of getting the answer wrong is high or the team genuinely lacks the prerequisite skill and the cost of failure is high. Pole B is right whenever the leader is acting from confident position on a question the team is closer to than the leader is, or whenever the leader is solving problems faster than the team can grow into.
In decisions: Pole A leaders open meetings with a position statement and shorten meetings by deciding faster; Pole B leaders open with “What’s on your mind?” and lengthen meetings by asking the five Toyota Kata questions, accepting that the team’s answer may be better than theirs.
Pole A. The leader is the most senior technical authority; problems are escalated up to where the expertise is.
Pole B. The leader’s role is to make others’ decisions possible; expertise lives at the work.
Pole A is right when the leader is genuinely the deepest expert and the call is technical: the founding CTO making the last call on a core architecture decision, the chief surgeon overruling on the table, the lead investigator on a specific case. Pole B is right when the leader’s expertise is a layer or two removed from where the work is actually being done, which is almost every senior leadership context outside narrow technical-founder situations.
In decisions: Pole A leaders are the technical decider on contested calls; Pole B leaders convene the people who know and facilitate the call they make.
Causation: how does this leader read the relationship between events and their causes?
23. Deviation vs conditions.
Pole A. When work goes wrong, the explanation is a deviation: from procedure, from competence, from intent. The fix clarifies the standard or addresses the deviator. “Human error” is a sufficient diagnosis.
Pole B. Real work continually adjusts to underspecified conditions; the gap between work-as-imagined and work-as-done is where both safety and risk live. Failure is the unexpected combination of normal variability; “human error” is a symptom of that gap, not a diagnosis.
Pole A is right in tractable Complicated/Ordered failures: a broken bearing, a software null-pointer exception, a structural calculation error, an individual whose behaviour was genuinely reckless. Pole B is right in any incident where the failure was the unexpected combination of normal variability, which is most of what goes wrong in Complex/Unordered work.
In decisions: Pole A leaders write performance-management documents, root-cause reports, and policy clarifications; Pole B leaders go to the gemba (the Japanese term for “the actual place”, where the work happens), ask “what surprised you? where did you have to improvise?”, and change the system people are operating within.
24. Safety as absence vs safety as presence. This is a corollary of entry 23, framed at the level of metrics and investment rather than incident analysis.
Pole A. Safety is the absence of events; the fewer accidents, incidents, and near-misses, the safer the system. Safety investment is insurance against accidents.
Pole B. Safety is the presence of defences, the ability to succeed under varying conditions. Safety investment is investment in productivity.
Pole A is right in tractable systems with stable specifications: the lost-time injury rate on a mature manufacturing line, the field-failure count for a long-standing physical product, the audit-finding count on a stable regulated process. Pole B is right in intractable socio-technical systems where work-as-done routinely adjusts to conditions work-as-imagined did not anticipate: a software platform under continuous deployment, a clinical service operating with staffing variability, anything where the absence-of-events metric stops moving while everyone reports getting away with more than they used to.
In decisions: Pole A leaders count injuries; Pole B leaders ask “how are we able to do this work well 999,999 times out of a million, and what do we need to preserve about that?”
See also. The sources under entry 23, Dekker’s Safety Differently and Allspaw’s How Your Systems Keep Running, both teach this dichotomy at the metrics-and-investment level too.
Pole A. Beginning-middle-end arcs with clear heroes, villains, and lessons.
Pole B. Several stories simultaneously about the same event; the discomfort of multiple frames as the data, not as a sign that the analysis is incomplete.
Pole A is right when the audience needs to act and the action is genuinely unambiguous: the all-hands message after a successful product launch, the board narrative on a clean quarter, the customer communication after a single root-cause incident with a documented fix. Pole B is right whenever a confident single-cause story is being constructed under conditions the data does not support: most strategic narratives, most multi-quarter performance explanations, most post-incident reviews where “human error” was offered as the answer in the first 24 hours.
In decisions: Pole A leaders narrate quarterly results as a story of decisions made; Pole B leaders narrate them as a story of conditions encountered, with several versions on offer.
Watch. Garvey-Berger, Mindtraps — Simple Stories (5 min). Companion: Tyler Cowen, Be suspicious of (simple) stories (TEDxMidAtlantic, 16 min): names the narrative fallacy directly and argues for “the mess” (multi-causal reality) as the alternative to single narrative. “Every time you tell yourself a good-vs-evil story, you’re lowering your IQ by 10 points.”
Self: how does this leader read their own identity, capacity, and developmental edge?
Every other dichotomy is mediated by the leader’s relationship to themselves. The developmental literature drawn on here (Kegan, Lahey, Garvey-Berger, Joiner) has decent measurement reliability for trained scorers, modest construct validity, and contested predictive validity. The dichotomies below name observable behavioural patterns rather than relying on staging interpretations to do the diagnostic work.
26. Socialised vs self-authoring vs self-transforming.
Pole A. Identity is given by the surround: role, profession, expectations.
Pole B. Identity is a self-chosen stance; the leader can hold their own values against pressure from valued others, and the role they hold is one expression of who they are, losable without the self being lost.
Pole C. Identity is held lightly; the leader values their own frame and watches for the data that would refute it.
Pole A is the right response in some contexts: early-career apprenticeship, certain professional-formation contexts, communities of practice where the work depends on shared identity. Pole B is what most senior leadership roles structurally require. Pole C is rare and not necessary for most roles, but is the right fit in the most VUCA situations where holding any paradigm too long is itself the risk: what Meadows called transcending paradigms, the highest leverage point in her hierarchy.
In decisions: Pole A leaders cannot disappoint valued authorities on behalf of their own judgment, and cannot imagine succession because the role is doing the identity-work; Pole B leaders can disappoint valued others, and treat succession as an explicit part of how they hold the role; Pole C leaders can disappoint themselves on behalf of evidence that contradicts their judgment. A three-pole entry; the gap between A and B is usually more salient than between B and C.
Identity-given-by-surround is the structural failure mode AI exposes most cleanly. Pole A leaders cannot distinguish good from plausible: the LLM’s confident prose lands in the same register as the voices that shaped their expectations. They accept what the model tells them to do not because they’re naïve but because they have no felt sense of what fitting would mean; the model’s output is just another voice from the surround. Pole B leaders feel the dissonance when a confident output doesn’t fit the situation, and slow down. Pole C leaders treat every model output as a hypothesis to be tested against evidence the model couldn’t have seen. The leader who cannot disappoint a confident voice cannot govern an AI either.
27. Subject to assumptions vs object of assumptions.
Pole A. The leader’s frame is what they look through; it is invisible to them.
Pole B. The leader can look at their own frame; they can hold it up to inspection.
Pole A is the structural default; every leader is functionally subject to most of their assumptions most of the time, and a working organisation depends on the leader operating from a frame rather than constantly inspecting one. Pole B is right where the situation is already telling the leader their frame is not fitting (results recurring, conversations repeating, the same diagnostic appearing in different vocabulary each quarter). The diagnostic move is the noticing in those moments, not the constant inspection.
In decisions: Pole A leaders are surprised by their own pattern when it surfaces; Pole B leaders can name the anxiety that drives the behaviours they keep noticing in themselves.
Watch. Kegan, An Evening with Robert Kegan and Immunity to Change (Boston College OD Network, 14 min). Opens with the 14-frogs-on-a-log puzzle (gap between deciding and doing) and names the immunity-to-change framework. Walks through identifying big assumptions, then lands on “being able to look at the whole thing”: the subject-object move on one’s own frame.
Pole A. The decision must yield this result; if it does not, the decision was wrong.
Pole B. The decision was a test of a hypothesis; the outcome is data; the question is what to learn from it.
Pole A is right when the decision was a genuine commitment to a specifiable outcome and the accountability frame requires defending the commitment. Pole B is right when the decision was a hypothesis dressed up as a commitment, or when defending the commitment is doing more work for the leader’s identity than for the outcome.
In decisions: Pole A leaders defend the decision after it fails; Pole B leaders refine the model after it fails.
Watch. Astro Teller, The unexpected benefit of celebrating failure (TED, 16 min). Kill-and-pivot stories from Alphabet’s X: vertical farming, cargo blimp, self-driving cars (the supervised model failed, the team switched to fully autonomous). “We bonused every single person on teams that ended their projects.” Closes on “enthusiastic skepticism is optimism’s perfect partner.”
The fifteen-minute version, named at the top, is enough on its own: look back at the six axes, notice where you sit on the left, and ask where those assumptions came from and whether they still fit. If the noticing surfaces something that wants more than a glance, the two-week version below is built for that. Four movements, four calendar slots, about two weeks end-to-end.
Movement 1, the reading pass. About 1 hour, today. Read the essay through once and flag the three or four entries that strike you on first pass as most provocative for your current situation. Then sample the embedded videos and essays under the Watch/read lines for those entries; most are 5-15 minutes, a few are longer. Notice which of your initial flags shift, drop, or sharpen once you’ve watched the videos or read the excerpts for those entries. The settled list is what you take into Movement 2. The point of this step is to keep this mirror itself from doing too much of the work; the sources are doing the teaching, and your first read should reflect what you actually find when you let them.
Movement 2, the dichotomy walk. 20 minutes alone, this week. Take your settled three or four entries into solo work. For each, ask three questions: which pole am I claiming?, which pole would my last three decisions in this territory actually show?, which pole does the situation I am facing actually require? The third is the Cynefin discipline: the situation determines the right pole, not the leader’s preference. The first two are the espoused/theory-in-use discipline (entry 11, the load-bearing wall). The interesting territory is wherever those three answers diverge. Write down what you find; you will need it for Movement 3.
Movement 3, the witness round. Send by Friday; one-week response window. Send the same three or four entries to three or four people who have watched you make recent decisions in the territory. Ask them which pole they observed in your last three decisions. One sentence each is enough. Compare to your own answer from Movement 2. The gap, when it appears, is the diagnostic.
Movement 4, the conversation. 45 minutes with a coach, the week after. Take the gap to a coach and work it: make visible the operating pattern that is already happening, and make the next decision a choice rather than a reflex. The achievable end state is one sentence you can say out loud: “On entry N, my espoused pole is X, my observed pole is Y, and the situation usually calls for Z.” If you do not yet have a coach who works at this depth, finding one is the work; this mirror is built for that conversation, and the conversation is what makes this mirror do anything other than sit on the page.
A staff-runnable variant. For a leadership team rather than an individual: same four movements, no coach. Movements 1 and 2 stay solo. Movement 3 is the team being each other’s witnesses on a single shared entry (entry 1 is a strong first choice). Movement 4 becomes a 30-minute conversation about the gaps that surfaced: what each person’s espoused pole was, what their last three decisions in this territory actually showed, and what the situation the team is facing actually requires. Run on one entry per fortnight. The team’s situation, not the facilitator’s, determines what each conversation produces.
What this mirror does, and what it doesn’t
This mirror makes the operating paradigm observable. It does not, on its own, change it. The change-work is the conversation that follows. The mirror is the first step of that work, not a substitute for it.
Three things you might do next.
If one entry hit hard: send it to one peer who has watched you make recent decisions in that territory, and ask them which pole they observed. One sentence is enough. That single exchange will tell you more than another read of the article will.
If you want to talk: flick me a note on LinkedIn, or visit hi.chrisgagne.com. If you do not yet have a consultant and/or coach who works at this depth, finding one is the work.
Some book links are Amazon affiliate links. If you buy through them I earn a small commission at no cost to you. Primary-source citations (papers, reports, articles) link to the original or to free archives.
Remote and hybrid work can take some getting used to. When you and your colleagues move from in-person to video-based collaboration, you miss out on some of the subtle body language that helps you know how to adjust your discussion flow.
Coaching globally distributed teams, I found these hand signals genuinely useful, so I thought I’d share what we learned. We used variations of them for years.
Before you start, make sure everyone can see everyone else at once. Most video platforms have a grid or gallery view that shows all participants on one screen.
The “stack” to facilitate taking turns speaking
The first set of hand signals is simply “the stack.” If someone else is speaking and you’d like to talk next, raise a finger. If someone already has a finger up and you’d like to speak after them, put two fingers up, and so on. I’ve seen stacks of up to five fingers.
When you’re done speaking, take a quick glance at everyones’ video feeds and call on the next person in the stack.
If you’re in the stack and still waiting when someone drops off the stack, lower one of your fingers so your place in line stays clear.
Try not to lower your hand for long, because that signals to the rest of your team that you no longer wish to speak. It’s polite to both go to the end of the stack if you still want to speak and let others get back in where they were if they try to return there.
The “C” to request a clarification with a question or comment
A “C” hand signal means you’ve got a clarifying question or comment. Anyone showing a “C” cuts to the front of the line and gets to speak before others in the stack. Use your judgement and try not to abuse this privilege.
If you’re speaking now when someone displays a C, call on them next. I rarely see more than one person throw a C at a time, but I’m sure you’ll find a way to work it out.
The “delta” to request a change in topic
Making a triangle with your two hands is called the delta. A delta, ∆, in math terms is a change. When you use the delta sign, you’re signaling that you’d like to see a change in topic or approach. If you’re speaking, try to finish soon and check in with the person displaying the delta.
Wiggling your fingers or thumbs up to express enthusiasm or agreement
If you are enthusiastic or agree with what someone is saying, you can put up your hands and wiggle your fingers with varying degrees of vigor. You could also show one or both thumbs up.
“Temperature check” to request feedback on a proposal
Sometimes you want to get an analog temperature check on how everyone’s feeling about something. To ask, just say something like “Let’s do a quick temperature check. This is ‘I’m hot on this.’ This is ‘I’m luke-warm on this.’ This is ‘I’m cold on this.” Then just give folks a few moments to respond accordingly.
“Roman voting” to gain consensus or feedback
You can also get a little more decisive with what Agile coaches call “Roman voting.” You can ask people to show you a thumbs up, a sideways thumb, or a thumbs down.
A thumbs up means that someone agrees with the proposal, a sideways thumb means they don’t necessarily agree but will go along with it, and a thumbs down means they disagree.
I suggest holding three working agreements around this:
If nobody is displaying thumbs down, you can move on. If you’ve got nothing but sideways thumbs, however, realize your proposal’s support is weak.
As the person seeking consensus or feedback, sincerely seek and consider the feedback of individuals displaying thumbs down.
If you show a thumbs down, be prepared to explain your objection and ideally what changes to the proposal could get you to a sideways or thumbs up.
To use Roman voting:
Make sure everyone knows your working agreement about what the hand signals mean and how you’ll use your feedback. For instance, will you proceed even if you get a thumbs down that you can’t resolve?
Ask people to raise a closed fist when they’ve made a decision and they’re ready to vote. Having people put up their closed fist first and voting all at the same time helps you reduce the effects of peer pressure and group think.
Vote on a quick count of three once everyone is ready.
Make a decision and move on, or discuss and vote again if necessary.
You can use these hand signals in person, too
I find that many of these hand signals are useful even when you’re meeting in person.
The “stack” and “C” hand signals help ensure that more people have a voice. Without them, people who are faster to respond—often the highest paid people in the room—often speak before others get a chance. By creating a stack, you help ensure everyone with something to contribute can be heard.
The “delta” is a way of suggesting that the meeting might be getting off-topic or digging into something too deeply for the time or audience. Teams who use a lot of sticky notes can also display a pad of stickies like a football (“soccer”) referee showing a “yellow card.”
Be mindful of your cultural context
I can’t think of too many hand signals that are both universally familiar and inoffensive. So it is important to choose signs that are appropriate for your cultural contexts. This is a balance of familiarity against potential offense: the “thumbs up” or “OK” signals are innocuous and familiar for the teams I supported, but both are offensive in several cultures those teams didn’t include. The specific sign is less important than the shared understanding of the intended concept, so work as a team to find familiar and inoffensive signs for your cultural context that represent the same concept.
In summary
You can use hand signals to augment your communication so that your team can work more effectively together. If you like these, share this article with your colleagues so you can come to a common vocabulary quickly.
What’s most important to your organization isn’t the adoption of these hand signals, but rather that you know you can add additional bandwidth to your communication using hand signals. You can always start with these and create new signals as part of your team’s Working Agreement or at your next retrospective.