You list five options, score them against criteria, and pick the highest scoring. The experienced operator already knew what option to choose, their skill long since dropped below conscious thought. You’re slower and you’re no more accurate. Nothing is wrong with the analytical method; it was simply aimed at the wrong decision.
Chapter 12 named the capacity to disagree with a confident voice; recognition is the cognitive-science version of the same question: when expert intuition is reliable, and when it is bias dressed as competence.
🎧 Prefer to listen? This chapter is narrated in my own voice with ElevenLabs on Spotify (24 minutes).
The two poles
Pole A: analysis. The right answer comes from listing the options, comparing them against criteria, and choosing the best. Deliberate analysis makes a decision defensible.
Pole B: recognition. The experienced operator recognises a situation they have seen before, retrieves one workable course of action, and mentally simulates it forward. It often arrives as a feeling before a thought, the gut that says roll back before you could say why. Pattern recognition makes a decision fast and—under the right conditions—reliable.
Where Pole A is right
Pole A is right when the environment is irregular, feedback is delayed or absent, the decision will need to be justified in public, or the operator lacks the experience that Pole B needs. Think of a CFO presenting capital allocation to the board, a new engineering manager making their first architecture call, or anyone facing a failure mode nobody has seen before.
Analysis is a good approach in those settings. The reader analysing a new acquisition, a new market, or a never-seen failure mode is doing the right work. For the engineering leader spending the day on AI governance, never-seen failure modes, and calls that will be justified in public, sustained analysis is the right choice. Recognition keeps the ground it has earned.
Honour where the analytic habit came from. A leader who came up justifying every call to a board, an audit committee, or a regulator built a real discipline, one that earns trust with other people’s money and works well for the decisions it was built to serve. The error comes when that discipline crosses into a domain where a calibrated expert’s first read is the better instrument and further thought buys delay and false confidence.
Where Pole B is right
Pole B is right when the environment is regular enough for experience to build trustworthy patterns and feedback arrives quickly enough to keep those patterns calibrated. Emergency response and surgical theatre fit. So does a deployment rollback that an expert has handled a thousand times, with the result visible within minutes.
In decisions
Pole A leaders use a checklist or formal review where the stakes are high and feedback is poor. Pole B leaders trust the experienced operator’s first reasonable option, while watching for the moment the situation slips beyond the operator’s experience.
The poles don’t carry equal weight here. Recognition owns a real but bounded territory: well-characterised, fast, recoverable judgements that an expert has made a thousand times. Analysis owns the rest, including the never-seen failure mode, the call the board will see, and the AI-governance decision that often sits behind both. Most of a senior leader’s hardest calls now sit there.
A CTO or VPE should carry one irony to the CEO on a single page: as AI takes over routine recognition, it erodes the very expertise operators need to catch its mistakes. The organisation’s approach to analysis and recognition decides how it governs AI tooling, and that governance belongs with the CEO.
The CEO needs to fund ways to preserve hard-won human expertise (deliberate practice, red-team exercises, and rotation of experienced operators through the edge cases AI handles least well) even when the dashboard is green. The budget won’t protect that work without a CEO-level signal. The sentence your CEO can carry to the board: “Automation is eroding the very expertise we’ll need on the day it fails.”
The Klein-Kahneman agreement
Gary Klein and Daniel Kahneman spent two decades disagreeing about whether expert intuition was real. Klein had documented it in firefighters, nurses, and chess masters. He called the mechanism Recognition-Primed Decision: the experienced operator recognises the situation as a type they have seen before, retrieves one workable course of action, and mentally simulates it forward (Klein, Sources of Power). Recognition makes the decision fast by replacing the comparison of options.
After Gary Klein, Sources of Power (1998); Klein’s own words: “the first workable option, not necessarily the best.”
Daniel Kahneman had spent the same decades synthesising evidence against reliable intuition in stock pickers, political forecasters, and clinical psychologists. Confidence in those settings tracked the coherence of the story rather than the quality of the evidence. On the surface, Klein and Kahneman couldn’t both be right.
They eventually agreed in a 2009 American Psychologist paper titled Conditions for Intuitive Expertise: A Failure to Disagree. Expert intuition is real where the environment is sufficiently regular to be predictable and the operator has had prolonged practice, with feedback, at learning its regularities. Outside those conditions, confidence is not evidence that intuition is reliable.
The conditions decide which pole fits, not the seniority of the person deciding. Gary Klein and Daniel Kahneman ended two decades of disagreement by naming when expert intuition is trustworthy: a sufficiently regular environment, learned through prolonged practice with feedback. After Daniel Kahneman and Gary Klein, “Conditions for Intuitive Expertise: A Failure to Disagree” (American Psychologist, 2009).
The conditions decide which pole fits. Ask whether they have produced reliable expertise.
Most organisations I have worked with ask whether the person making the call is senior enough. The Klein-Kahneman agreement asks whether the conditions have calibrated their recognition.
I recalibrate my own mentoring on this: the same misfire turns up in clients whose recognition was built in a context that has since changed.
Klein’s method underneath the framework
Klein reached Recognition-Primed Decision through fieldwork. His Critical Decision Method, a form of cognitive task analysis, asks experienced operators about recent hard incidents. The interviewer follows the story and probes the thinking: what did you see?, what made this familiar?, what would someone with less experience have done?
In regular environments, experienced operators usually retrieved one option, simulated it, and acted. They didn’t generate a set of alternatives. Herbert Simon had named satisficing decades earlier as a cognitive limit: taking the first option that meets the relevant criteria rather than maximising across all of them. Klein showed when that limit becomes an achievement: in regular environments, under time pressure, the expert takes the first workable option and acts.
A senior site-reliability engineer who calls a rollback in 15 seconds may be doing exactly what experienced operators do in regular environments with rapid feedback. A manager who makes them stop and list five options imposes Pole A on a Pole B situation, slowing the decision without improving it.
Distributed cognition
Recognition reaches beyond the individual. Cognition is distributed across people, tools, and the environment. A ship’s navigation team is one cognitive system, with the navigator working as part of a wider whole. When something fails on the bridge, the configuration that brings the ship to anchor includes the charts and instruments, the bearing books and trained crew, and the procedure for fixing position.
Remove just one component and the system can no longer think in quite the same way.
Engineering organisations also hold expertise across a system. The senior engineer who knows the legacy code depends on source code, a build pipeline, and an on-call playbook; none carries the whole capability alone. Pole A leaders sometimes ask individual engineers to compensate for a poor configuration through more analysis. Pole B leaders improve the configuration itself with better tooling, documentation, and playbooks, so the system carries more of the work.
The safety-science chapters that follow make the same mechanism concrete. The operator is part of the system, so blaming them treats one component as though it were the whole. Chapter 14 takes that claim out of cognitive science and into safety.
Bainbridge’s ironies of automation
Lisanne Bainbridge wrote a five-page paper in 1983, Ironies of Automation, that names the central mechanism behind the recognition-versus-analysis choice in AI-augmented systems. Forty years before AI agents, she named the cost of letting the machine do the easy parts.
Bainbridge named two ironies. First, the designer who tries to eliminate the human operator is also a potential source of system failure. The second deserves more time. As Bainbridge put it, the designer who sets out to remove the operator still hands that operator whatever the designer couldn’t work out how to automate, so the operator ends up with a leftover, arbitrary set of tasks that nobody designed support for. Automation strips the easy parts and leaves the hard ones. Then it gives the operator little help with what remains.
Automation strips the easy parts and leaves the hard ones. Then it gives the operator little help with what remains.
Bainbridge catalogued the effects. Physical skills fade without use, so a formerly experienced operator who has spent long enough monitoring an automated process may now be an inexperienced one. Cognitive skills also need frequent use and feedback; once automated, they lose the patterns recognition depends on.
Working storage is Bainbridge’s term for the operator’s running feel for the process, richer than short-term memory, and it takes time to build. She noted that some manual operators enter a control room a quarter to half an hour before they are due to take over control so they can recover that feel. The central irony remains: the automatic control system is installed precisely because it outperforms the operator, and yet that same operator is the one charged with monitoring whether it is working.
Her verdict, in her words: “The human monitor has been given an impossible task.” The operator must catch what the automation misses in a domain where the automation already outperforms them on routine cases. Meanwhile, the automation has hollowed out the expertise built under the conditions Klein and Kahneman agreed produce reliable judgement.
The human monitor has been given an impossible task.
Lisanne Bainbridge, Ironies of Automation, 1983
Bainbridge’s irony of manual takeover follows the same mechanism. By the time a human has to take over, something has usually gone wrong with the process, so controlling it calls for unusual action, and you can argue the operator at that moment needs to be more skilled and less loaded than average, not less skilled and more loaded.
At the moment of need, the operator is less skilled and more loaded than they were before automation took the routine work.
My read is that Bainbridge’s 1983 mechanism explains AI governance better than most contemporary work: automation weakens the expertise it still depends on. The problem is older than ChatGPT. The question is whether we have any better answers than we did in 1983 and—from the evidence I’ve seen to date—we mostly don’t.
Pole B recognition works only when the operator has the orientation the situation needs. Automation that strips away the operator’s experience degrades the conditions that make Pole B trustworthy. The longer the system runs well under automation, the less reliable the operator’s recognition becomes, until the moment something goes wrong and the system needs that recognition most. Chapter 17 returns to the authority problem this creates.
Automation is eroding the very expertise we’ll need on the day it fails. The lines draw the shape of Bainbridge’s argument, not data. After Lisanne Bainbridge, “Ironies of Automation” (1983); “impossible task” is her verbatim verdict: “The human monitor has been given an impossible task.”
The AI failure modes Bainbridge is pointing at
Frequency-versus-truth is the first failure mode. A large language model predicts the next token from statistical patterns in training data. Its output tracks the distribution of that data more than the truth of any claim. The training distribution is dominated by the average of what humans have written. That pulls the model toward the average consultant’s answer, the regression chapter 1 described, not the contrarian signal that makes a diagnosis valuable.
The model reproduces common errors with confidence because confidence tracks coherence of story, not quality of evidence. This is the Kahneman side of the Klein-Kahneman agreement applied to the system itself. Chapter 1’s grounding argument is the response, though the leader can’t always tell from the output whether grounding is in play.
Out-of-distribution collapse is the second failure mode. AI models perform well on tasks that resemble their training data. They degrade, sometimes gradually and sometimes abruptly, on tasks outside it. The model’s confidence doesn’t track its accuracy across this boundary; the output looks the same whether the task is well within the training distribution or well outside it.
The model performs well on one task and badly on a task that looks almost identical, with nothing on the surface marking the difference. It produces plausible-looking code outside its distribution, the code fails contract tests the model was never given, and the operator ships it on trust without running the check.
Bainbridge expertise-hollowing is the third failure mode. An engineer who hands straightforward problems to the model keeps responsibility for the hard ones, exactly as in Bainbridge’s manual-takeover scenarios. At the same time, the engineer loses the practice needed to handle those hard problems: the pattern recognition built by solving easy cases and the feel for normal that makes abnormal recognisable.
The irony holds exactly. The automatic system was put in because it performs better on routine tasks, but the engineer must still catch the cases where it fails. At the moment of need, the engineer is less skilled than they were before handing the routine work over.
A worked example: the experienced incident commander
A whole group of users logged on one morning and couldn’t do their jobs. Within two minutes of the bug being raised, a senior engineer named the cause: duplicate user records, each person now holding two accounts, only one of them wired into the workflow they needed. They tested it on a single account first, deleting one duplicate and checking that person could work again. They could. The rest were cleared inside twenty minutes. They didn’t list five hypotheses and score them. They recognised the pattern.
Recognition is doing the work analysis would otherwise have to do. It is faster and, under the right conditions, more accurate because the engineer had seen this kind of failure before, the dashboard provides rapid feedback, and the diagnostic action is reversible if the hypothesis is wrong.
Recognition is doing the work analysis would otherwise have to do. The mechanism is Gary Klein’s Recognition-Primed Decision, Sources of Power (1998); the incident is from the chapter.
Take away any of those conditions and Pole A becomes the right move. A newer engineer may lack the pattern. An under-instrumented dashboard can’t provide quick feedback. An irreversible diagnostic action makes a wrong first read too costly. In those cases, the Pole A leader is right to ask for a checklist. Where the conditions support recognition, the Pole B leader is right to let the engineer decide.
Recognition can also point confidently at the wrong layer. In another incident a whole floor lost a screen they needed, and the failure read as infrastructure: about five hours went into the cluster, the pods, and DNS. The cause was a content-filtering policy on the office network blocking the domain. The wrong read held because the monitoring couldn’t see the problem: the health checks ran from inside the cluster, so they never crossed the firewall doing the blocking. The environment wasn’t supplying the feedback that makes recognition trustworthy.
The AI overlay
AI agents work like recognition machines: they identify patterns in training data and produce outputs that match. They are more dependable when the work resembles their training examples and the output can be checked quickly, which is where Pole B is right. They are less dependable when either condition fails, which is where Pole A is right.
Deploying AI in Pole A territory and trusting its confident outputs is a structural mistake. The model has no patterns to recognise there, so coherence of story masquerades as expertise. The Kahneman side of the Klein-Kahneman agreement applies to models as it does to people. Confident output without environmental regularity is bias.
Bainbridge’s ironies follow. When AI handles routine recognition well, the human operator loses the practice needed to catch the cases it gets wrong. Governing AI therefore calls for Pole A discipline: keeping human expertise alive while automation does most of the routine work. Chapter 17 returns to this AI-and-authority question.
The recognition story still holds within a smaller domain than the one a senior leader now inhabits. Once AI is in the loop, the human must keep a live, independent read on the automation because its confident output is least trustworthy in irregular territory. Sustained analysis is the governing discipline.
A leader governing AI well first identifies which decisions fall in regular environments, where AI recognition is reliable, and gets out of the way. For irregular environments, the leader uses analysis, checklists, and formal review; most consequential calls belong here. The leader also preserves the expertise automation would otherwise erode through deliberate practice, red-team exercises, and rotation through edge cases, so the human keeps an independent read on whether the automation is wrong.
Two of those three are Pole A work: formal review for the irregular calls, and the deliberate practice that keeps human judgement sharp. Applying one method to all of them is bad governance, so the discipline is routing by conditions rather than a standing default. For the AI-governing reader, those conditions point to analysis often enough that the chapter resolves there.
The diagnostic move
Three questions for last Tuesday’s decision:
Which pole was I claiming? Did I treat the decision as one for analysis or for recognition?
Which pole would the actual decision-process show? If a colleague watched the process, would they see deliberate option-comparison or experienced-operator recognition?
Which pole did the conditions actually support? Regular environment, rapid feedback → Pole B. Irregular environment, delayed feedback → Pole A. Automation-mediated environment that has hollowed out the operator’s experience → Pole A by structural necessity, regardless of what the operator’s stated experience suggests. Automation-mediated calls are becoming common enough that a leader whose work is increasingly AI-mediated should test for hollowed-out experience explicitly, which is why the chapter resolves toward analysis for that reader.
In my experience, the most common gap is between the second and third questions. Leaders trust recognition in conditions that don’t support it, or apply formal analysis where the experienced operator already knew.
The exercise
Run a Klein-Kahneman conditions check on a decision your team is sitting with this week. Ask two questions aloud. Is the environment regular: have you and your team seen many cases like this? Is feedback rapid: will you know within days whether the call worked?
If both answers are yes, trust the experienced operator’s recognition and get out of the way. If either answer is no, run an analytic protocol.
Then try the check on a decision the team has already made and is now defending. Where the check turns up a no answer on a defended decision, the operator’s confidence outran the conditions. Naming that without blame builds the check into the team’s practice.
In-text: the central pair and the AI overlay. Gary Klein, Sources of Power; Daniel Kahneman, Thinking, Fast and Slow, on the Kahneman-Klein agreement (the 2009 American Psychologist paper Conditions for Intuitive Expertise: A Failure to Disagree). For the AI-automation irony 40 years early: Lisanne Bainbridge, Ironies of Automation (1983; 2021 retrospective preface available online).
Go deeper: the same recognition claim arrives from several further traditions, none anchored in the body above. For the OODA loop and orientation as the central node: Robert Coram, Boyd: The Fighter Pilot Who Changed the Art of War (and Boyd’s appendix essay Destruction and Creation; Fingerspitzengefühl, fingertip-feel). For the doctrinal statement of Pole B: USMC, MCDP-1 Warfighting, Chapter 4. For cognition distributed across people, tools, and environment (the USS Palau navigation case): Edwin Hutchins, Cognition in the Wild; and for joint cognitive systems as the unit of analysis: David Woods and Erik Hollnagel, Joint Cognitive Systems. For the wider bounded-rationality lineage the chapter inherits: James March, A Primer on Decision Making. And from the emotional-intelligence tradition, for the felt and bodily side of recognition a purely analytic read can’t reach: Karla McLaren, The Language of Emotions, on the four intelligences (analytical, emotional, somatic, visionary) and the cost of letting the intellect try to rule without the other three.
Bainbridge’s irony doesn’t stop at the console. Automation wears away the operators it leans on. One of them eventually slips, and you open the incident review to name the rule-breaker who caused it.
You find one.
The deviator was doing what the system trained them to do, so the report you sign punishes a person and leaves the conditions loaded for whoever fills the role next. Chapter 14 pulls the punishment off the person and puts it back on the conditions.
I work with engineering leaders on exactly this kind of paradigm work, the deeper the better. If it’s live for you, I’m happy to talk: schedule a 30-minute virtual coffee at hi.chrisgagne.com.
Some book links here are Amazon affiliate links; if you buy through them I may earn a small commission, at no cost to you.
Continue reading:Chapter 14, Deviation vs conditions.
Think of a physiotherapist. They don’t fix your back. They put a mirror in front of you, show you the exact posture you’ve been holding for fifteen years, and say: “notice you’re doing that? that’s why you’re in pain.” The pattern was invisible from the inside. Once it’s visible, the work is small corrections, practised many times, especially under load. This piece is that kind of mirror.
NB. The dichotomies below are mostly not original to me; the curation and the synthesis are. I name teachers in each entry and link to their work so you can go upstream when something resonates. If this mirror surfaces something worth working, the people whose names appear in the Watch and Read lines are where I’d send you next. Or me—this is the work I do.
When senior leaders see the operating paradigms they have been running on, the calls they make under pressure stop surprising them. With that, the response that better fits the new situation becomes available. The work I do with leaders suggests most are running paradigms they cannot see, built for problems the firm no longer has, in territory the paradigms weren’t built for. AI is the most visible driver of the changing context, but not the only one. The piece below offers a mirror for seeing what’s running underneath: 28 observable decision-level dichotomies, written as prompts to compare what you say you believe to what you actually do.
This is the canonical long-form treatment of the work; a shorter, axis-by-axis article series will follow for readers who want to take it in smaller doses.
Who this is for, who it isn’t. This is for senior leaders of software-dependent firms whose situation-mix has shifted toward complex, unpredictable work: long-term survival under uncertainty, AI-era operating-model threats, leading technology-driven market disruption. You might suspect the mindset that got you where you wanted to be yesterday isn’t fitting what your new context needs today, or you might find yourself in the midst of a crisis. It isn’t a leadership-style typology, a developmental ladder, or a self-administered test. You are not broken and do not need fixing; you are simply running on paradigms that fit a context you have since left. If your work is genuinely and predominantly Complicated/Ordered (regulated compliance, financial close, structural engineering, mature manufacturing on a stable specification), the classical management paradigms are possibly sufficient. That said, almost no senior role in a software-dependent firm is purely Complicated/Ordered anymore: take thirty minutes anyway, and notice how much the circumstances have shifted underneath you.
The diagnostic question. Most knowledge work has been Complex/Unordered for decades: the territory where the answer isn’t knowable in advance, only discoverable by probing (Dave Snowden’s Cynefin, defined more fully in “The frame” below). What changed is the cost of mismatch. The Complicated/Ordered habits—the classical management toolkit, built for work where experts can analyse the problem to a known good answer—could carry a slow-moving firm through a Complex/Unordered situation as long as the consequences of being slightly wrong arrived slowly too. With AI accelerating the consequences and reshaping the substrate, the gap between the leader’s habitual response and what the situation actually requires now compounds in months rather than years. So the diagnostic question is: my role has been Complex/Unordered all along; is my habitual response still fitting now that the cost of mismatch is, for a growing number of firms, terminal?
On Chegg’s May 2023 earnings call, CEO Dan Rosensweig named ChatGPT as the threat to new-customer growth; they launched CheggMate, their OpenAI-built tutoring product, the same month. Three years later the market cap is down roughly 99% from its 2021 peak, headcount has fallen from ~3,200 to under 600, and the Q1 2026 homework business is down 57% year-on-year. They posted their first profit in two years by cutting costs faster than revenue fell. Awareness was not the limiting factor; the org could not redesign around the substrate shift in time.
The three most important paradigm dichotomies for most readers in 2026. If you read nothing else, read these:
Entry 1, AI as superficial productivity layer vs AI as substrate for org redesign. The dichotomy that decides whether the firm’s AI investment compounds or runs out as a local productivity uplift that doesn’t translate to revenue.
Entry 6, Utilisation vs flow. The dichotomy that decides whether AI speeds your system up or just relocates the queue.
Entry 11, Espoused theory vs theory-in-use. The dichotomy that decides whether your stated principles survive contact with your last three decisions under pressure.
The six paradigm axes and the two poles. Each axis asks where your operating reflex sits. Pole A is the move that often feels like leadership; Pole B is what the Complex/Unordered contexts reward most. Both poles have bounded applicability, neither is correct for all contexts. The question is whether your habitual response fits the situations you actually meet.
Change: how this leader produces change.
1. AI bolted onto the existing organisational design, or the organisational design entirely reconsidered around AI.
2. Announce the change programme and run training, or redesign what surrounds the behaviour so the new move is the easier move.
3. Write the strategy and roll it out, or design the conversation that produces the strategy.
4. Ship a directive, or ship a question and an open week.
5. Send people on courses, or block thirty minutes a day for coaching at the work.
Optimisation: how this leader reads the unit of value and the shape of waste.
6. Keep everyone fully utilised, or keep queues and lead times short.
7. Penalise plan variance, or plan a hypothesis and an experiment.
8. Set efficiency targets per function, or find the one constraint and subordinate everything to it.
9. Celebrate the launch, or celebrate the customer succeeding three months later.
10. Long roadmaps and/or Planning Intervals, or continuous flow, WIP limits, and pull.
Knowledge: what this leader treats as legitimate knowledge.
11. Judge yourself by your intentions, or judge yourself by your last three decisions under pressure.
12. Project certainty, or voice a hypothesis and the test that would falsify you.
13. Run the checklist, or trust the experienced operator.
14. Standardise across business units, or protect what the local team already knows how to do.
15. Convene more analysis when stuck, or pause and check what the experienced operator already senses.
Authority: how this leader reads power, decision rights, and their own role.
16. Fill coordination channels with approvals, or fill them with “I intend to…”
17. Decide where the formal power is, or decide where the tacit knowledge actually lives.
18. Get the work out of people, or get out of people’s way.
19. Be the hero, or build up your successors.
20. Cascade the OKRs down the org, or redesign decision rights so the org can move without you.
21. Open with the answer, or open with a question you don’t already know the answer to.
22. Be the deepest expert in the room, or convene the people who know.
Causation: how this leader reads the relationship between events and their causes.
23. Find the deviator who failed, or close the gap between the work as imagined and the work as done.
24. Count incidents, or study what works well (not just what broke).
25. Tell one clean story, or hold several stories at once.
Self: how this leader reads their own identity, capacity, and developmental edge.
26. Hold the role as who you are, or hold it as what you’re currently doing.
27. Be surprised by your own pattern when it surfaces, or name the anxiety driving it.
28. Defend the decision after it fails, or refine your mental models after they fail.
15 minutes today. Look back at the six axes above. For each axis, notice which pole you sit closer to. Where you sit on the first pole, ask: where did this assumption come from, and is it still fitting the work you actually do now? That alone will be enough to start the noticing the rest of the article is built around.
If the 15-minute pass surfaces something worth more than a glance, continue reading and work the two-week protocol in How to use this mirror at the end.
Why a mirror, and why now
Jennifer Garvey-Berger puts it plainly: “We get so used to our patterns that we can forget what the situation calls for and just rely on our own habits.” That is the diagnostic problem this mirror addresses.
A senior leader who has read this far is still left with a practical question. Which paradigm am I actually running? The intellectual answer is almost always the flattering one. The honest answer lives in last Tuesday’s meeting, in the calls I made under pressure when nobody was watching me make them.
This mirror is ultimately meant to be used in conversation: with a coach, a trusted peer, a few witnesses willing to compare what you say to what they actually saw you do.
It shows you what habit has hidden from you. The change is the small correction—repeated under load—in the conversations that follow. One mirror is rarely enough; like a physiotherapist’s, the role is distributed across several people who will keep showing you the same pattern from slightly different angles until you can feel it yourself. You’ll need at least a consultant who understands the paradigms and a developmental coach who can help you through the change; these are rarely the same person.
The frame: the situation-mix in your role has shifted
Snowden’s Cynefin distinguishes five kinds of situation:
Clear. Best practice, fixed constraints, self-evident cause and effect.
Complicated. Good practice via expert analysis, knowable causal relationships, governing constraints.
Complex. Exaptive practice via probing, enabling constraints, emergent and dispositional behaviour (which is interestingly something an LLM also embodies).
Chaotic. Novel practice, no effective constraints.
Confused. Knowing which domain you are in is itself the question.
Cynefin is a sense-making framework for assessing situations, not for typing leaders. The same leader routinely meets situations in all of those domains and is asked to respond to each.
The diagnostic question this piece works with is therefore not “which domain do you operate in.” It is: has the situation-mix in your role shifted toward Complex/Unordered work, and does your habitual sense-making mode fit the situations you now meet?
For most senior leaders in software-permeated firms, the answer to the first half is plainly yes. Strategy under genuine uncertainty, leadership transitions, cross-functional change, AI-era organisation redesign, talent dynamics, market position under technological disruption: these are all Complex/Unordered. This has been true for most knowledge workers for decades. The classical management toolkit was built for Complicated/Ordered work—regulated compliance, financial close, structural engineering, mature manufacturing—and remains right in that domain. The trouble starts when a Complicated/Ordered habit meets a Complex/Unordered situation and treats it as if the answer were knowable.
Ronald Heifetz puts the same diagnostic in a different vocabulary:
Technical problems have known solutions accessible through current authority and expertise: forecast, plan, hold the plan as the basis for accountability; when results slip, take more direct control or buy the answer.
Adaptive problems require changes in values, beliefs, roles, or relationships, and no single action controls the outcome: run safe-to-fail experiments, hold the group in productive disequilibrium, let pattern emerge.
Heifetz’s adaptive-vs-technical and Snowden’s Complex/Unordered-vs-Complicated/Ordered are similar diagnostics at different scales; every entry below is a slice of the larger question: what does this leader do when the response that used to work no longer fits the situation?
I learned this exercise from Bas Vodde. Draw a horizontal axis. Theory X to the left (workers cannot be trusted, require coercion and control), Theory Y to the right (work can be as natural as play, people seek responsibility and meaning). Sticky-notes above the line are people’s observations of their colleagues’ attitudes and behaviours. Sticky-notes below the line are observations of the organisational context they’re operating inside. The gap between the two rows of sticky-notes, when it appears, is the data:
A redacted artefact from a Theory X / Theory Y exercise I facilitated. The top row clusters Theory Y; the bottom row clusters Theory X.
Pole A, Theory X. Design control structures around the worst-case employee; oversight as the default; trust as a reward earned through visible compliance.
Pole B, Theory Y. Design control structures around the average employee; trust the system to handle the exceptions; remove obstacles and let people do good work.
Where Pole A is right. Genuinely high-stakes regulated work where the cost of a single bad actor is catastrophic: nuclear plant control rooms, aviation cockpit checklists, surgical theatre protocols, anti-money-laundering compliance. The structural Theory X is not a failure of trust; it is the right engineering response to a Clear- or Complicated-domain risk surface in a low-variance context.
Where Pole B is right. Most creative work, most knowledge work, genuinely Complex cross-functional decisions, almost any role where intrinsic motivation is doing the real work.
What it looks like in decisions. Pole A leaders write detailed approval workflows for spending below $5k. Pole B leaders set a budget and let the team allocate.
Source. Douglas McGregor.
The larger mirror
Here are all 28 paradigm dichotomies, sorted into six roughly-organising axes.
Change: how does this leader read and produce change?
1. AI as productivity layer vs AI as substrate for org redesign.
Pole A. AI is a tool that makes existing roles faster, cheaper, or more leveraged. The organisation chart, role boundaries, and career ladders stay; capable people get more done. Pole A treats AI as adoption.
Pole B. AI is a substrate that re-shapes which work is human and which is not. The valuable human role broadens rather than narrows: the builder, or skilful generalist, who can define an outcome, specify the constraints, choose which work goes to AI and which stays human, and verify what comes back. The narrow specialist’s economic logic—expertise is scarce and demand for it is unlimited—is the assumption AI dissolves. Pole B treats AI as redesign and treats the generalist as the new load-bearing role.
Pole A is right where the work is genuinely procedural, the regulatory or domain constraints prevent role redesign within the planning horizon, or the firm is buying time to learn what to redesign. Pole B is right wherever the planning horizon is shorter than the rate of capability change in the AI substrate, or where competitors are already redesigning.
In decisions: Pole A leaders sponsor AI-rollout programmes inside existing job families and ask “what’s the productivity uplift?”; Pole B leaders ask “which work is now human, which is now AI-supervised, which is now AI-autonomous, and what does the role—if any—on the other side actually look like?” They accept that the answer reshapes the entire organisation’s strategy, process, people, structure, and rewards. The traps on both sides are different. Pole A leaders may well discover two years late that their biggest cost curve was the incumbent org design itself, not tokens. Pole B leaders get the org design right (adaptive topology, broadened mandates, smaller teams of builders) but might bet specific designs on AI capabilities that aren’t yet stable, and have to walk those back.
The builders carrying the substrate shift in the firms I see doing it well tend to share a profile: high pattern-recognition across domains, liberal-arts-educated as much as they are technically trained, comfortable holding business, technical, and human-systems vocabulary in the same sentence. They are noticeably less patient with the status quo than average. The org-design literature names the role (multi-learning, M-shaped, expert generalist) but undersells how specific the profile is. Hiring for it is harder than hiring for the role it replaces.
Watch. Craig Larman, Craig Larman’s Take on 10X ORG: Why He Thinks This Book Matters (5 min). Names the dichotomy directly: Pole A as the cost-curve trap (“AIs cheap as dirt, 24/7, no vacation, humans won’t compete by suggesting ‘hire me, I’ll make you 10% better’”), Pole B as 10x redesign rather than 10% improvement, with 10xOrg, Org Topologies, and LeSS. Companion: the Reinertsen video from entry 6 for the flow-economics half.
Read. Org Topologies (start with their primer) and 10xOrg on org-design at AI-substrate scale; Jay Galbraith, Designing Organizations (Star Model) on the model that underpins much of Org Topologies and 10xOrg; and Reinertsen, The Principles of Product Development Flow, on what happens to flow and cycle time when the speed of individual process steps changes faster than the surrounding system.
2. Pushing people vs changing the conditions.
Pole A. Produce change by force of will: announce the programme, run the training, hold people accountable for the new behaviour.
Pole B. Produce change by shifting what surrounds the behaviour: the path (incentives, friction, information flow) and/or the operating system (the meaning-making that produces the behaviour in the first place). Pole B has two flavours: environmental design (change the situation, behaviour follows) and developmental work (change the mindset, behaviour becomes durable). Change the organisation design and the behaviours will also change.
Pole A is right in time-critical regulatory or competitive shifts where the new behaviour is specifiable and adoption speed is the binding constraint. Pole B is right whenever the change is adaptive rather than technical, or whenever Pole A compliance reliably decays the moment pressure relaxes.
In decisions: Pole A leaders announce programmes and run training; Pole B leaders redesign the environment so the new behaviour is the easier behaviour, and accept that the deeper version of the work is developmental, takes much longer than scheduled, and treats early relapse as “not yet” (in Heath & Heath’s frame from Switch, drawing on Carol Dweck) rather than as failure. Training is still valuable—especially when it transmits difficult-to-document embodied wisdom—just not sufficient.
Pole A. Get the right answer; convince others; the decision itself is the work.
Pole B. Get the right way of deciding; the quality of the process determines the quality and durability of the decision.
Pole A is right when the answer is genuinely knowable and the legitimacy of the decision is not at stake: the technical call on a deployment rollback, the regulatory interpretation, the budget reallocation inside an already-agreed envelope. Pole B is right when the decision will be implemented by people who weren’t in the room and whose ownership of the decision is what makes it stick, which is most systemic decisions.
In decisions: Pole A leaders write the strategy and roll it out; Pole B leaders design the conversation that produces the strategy.
Watch. Roger Schwarz, Smart Leaders, Smarter Teams (4 min): the leader as architect of how the team thinks together, not just decider of what gets decided.
Pole A. Authority issues the requirement; people comply.
Pole B. Authority articulates purpose; people enrol or commit; compliance is the floor, not the goal.
Pole A is right when speed is binding and the cost of partial enrolment is acute. Pole B is right when the work depends on enrolment for any durability.
In decisions: Pole A leaders ship a directive; Pole B leaders ship a question and an open week.
Watch. Peter Senge, On Shared Vision (2 min): on the compliance-to-commitment ladder and why a vision has to be talked about with people, not told to them. Companion: Frederic Laloux, 1.2 What truly drives you? (Thought for top leaders) (10 min) on speaking from the why rather than mandating the what. The University Hospital CEO story shows the exact moment a leader stops mandating concepts and resistance vanishes.
Pole A. Episodic, off-site, formal; the classroom is the venue.
Pole B. Daily, embedded in work, mentor-mentee at the gemba; the workplace is the venue.
Pole A is right when the content is genuinely transferable from classroom to context and the transfer overhead is low: compliance certifications, language fluency, formal credentialing requirements, software platform onboarding for stable tools. Pole B is right when the content lives in the situated work and the classroom abstraction loses most of what matters: coaching skills, judgment under uncertainty, leadership behaviour under threat, and almost any capability the leader cares about in their senior people.
In decisions: Pole A leaders send people to courses; Pole B leaders block thirty minutes a day for coaching cycles at the work.
The exception worth naming: well-designed games, simulations, and case-based kata are training in venue but Pole B in substance. They transmit the tacit pattern-library—the mētis—that “sage on the stage” classroom abstraction usually can’t carry. Donald Schön’s practicum, Gary Klein’s recognition-primed decision simulation work, and Ikujiro Nonaka’s socialisation mode in SECI all name this third space. The diagnostic isn’t the venue; it’s whether the content is codified or situated.
Optimisation: what is the unit of value, and how does this leader read waste?
6. Utilisation vs flow.
Pole A. Keep resources busy; idle capacity is waste; efficiency reports are the management metric; manage cost first.
Pole B. Throughput is the goal; queues are the enemy; capacity margin is a feature, not a bug. As utilisation climbs toward 100%, queue time dominates cycle time and small upsets in arrivals cascade into long delays.
Pole A is right in genuinely fixed-throughput operations with stable demand and homogeneous work: call centres on routine queries, very mature manufacturing lines. Pole B is right in any system with variable demand, heterogeneous work, or genuinely }Complex/Unordered prioritisation.
In decisions: Pole A leaders allocate every engineer to a project and ask “are you fully booked?”; Pole B leaders protect slack, accept idle time at non-constraints, and ask “what are my queue lengths” as a leading indicator of lead times.
The three curves say the same thing: the optimum is not the peak. Utilisation as the metric runs you past it. Queue size is the textbook single-server curve, ρ/(1−ρ); the other two charts show shape only. After Don Reinertsen, FlowCon 2014.
Pole A. Strategy as document, vision as destination, plan as commitment; decide early, execute against the commitment; variance is failure.
Pole B. Strategy as verb; set direction, find adjacent possibles; treat each plan as a hypothesis with an expected outcome; check; act on the difference; “cross the river by feeling the stones.”
Pole A is right when the planning horizon is shorter than the rate of change, the cost of changing the plan is high, and accountability requires commitment. Pole B is right when none of those is true.
In decisions: Pole A leaders penalise plan variance and demand a guarantee; Pole B leaders demand a hypothesis and an experiment, update the plan, and treat the update as evidence the strategy is working.
Watch. Mary Poppendieck, The Tyranny of “The Plan” (60 min, transcript also available on LinkedIn; this is canonical and well worth the time).
Pole A. Improve each department on its own metrics; the whole will improve as a sum.
Pole B. Identify the one constraint that limits the whole, exploit it, subordinate everything else to it. “A system of local optimums is not an optimum system; it is a very inefficient system.”
Pole A is right only when the system genuinely decomposes: independent product lines, separate geographies, decoupled value streams with no shared constrained resource. But most systems leaders think decompose don’t. The resources, the people, the queue capacity, the management attention all run through shared bottlenecks the org chart hides. In manufacturing—a Complicated/Ordered domain—Goldratt demonstrates that local efficiency improvements on non-bottleneck resources reduced plant throughput, by pushing more work onto the bottleneck and consuming time that couldn’t be recovered. W. Edwards Deming’s funnel experiment proves the same point for stable production processes: tampering with common-cause variation makes the variance worse, not better. The claim isn’t restricted to Complex/Unordered work; it holds anywhere events depend on prior events and demand has variability, which is nearly every system a senior leader actually meets.
In decisions: Pole A leaders set utilisation targets per function; Pole B leaders ask “which one resource, today, is governing how fast the whole system can move?” and protect it.
Pole A. The unit of management is units shipped, features released, work completed.
Pole B. The unit of management is customer progress, market position, mission impact.
Pole A is right when output and outcome have been demonstrated to track each other tightly in the relevant domain: a sales organisation on a mature product where deals booked do convert to revenue, a fulfilment operation where units shipped equals customer demand met, a regulatory pipeline where filings submitted equals approvals secured. Pole B is right whenever the firm is producing more output and getting worse outcomes, a common pattern in feature-factory product organisations and in any system where the output metric has become the goal rather than the proxy.
In decisions: Pole A leaders celebrate the launch; Pole B leaders celebrate the customer using the thing successfully three months later.
Pole A. Aggregate work into larger units to amortise setup costs; schedule and dispatch work to resources based on forecast; central plan governs.
Pole B. Reduce batch size; resources pull work when they have capacity; demand signal at the customer end governs upstream.
Pole A is right when setup costs are genuinely fixed and high, forecast accuracy is high, and the cost of inventory is low (classical EOQ logic in mature physical operations). Pole B is right in any context where setup costs can themselves be reduced and the learning rate matters, or where forecast accuracy is low.
In decisions: Pole A leaders schedule long roadmaps and allocate capacity by Planning Interval; Pole B leaders ship continuously, set WIP limits, and let teams pull from a ruthlessly prioritised and pruned backlog.
Knowledge: what does this leader treat as legitimate knowledge?
11. Espoused theory vs theory-in-use (Model I vs Model II).
Pole A. The leader’s stated principles, values, and intentions are taken as the relevant data; when results miss, change the action while leaving the underlying assumptions intact. Under threat or embarrassment, Model I governs: win not lose, be rational, avoid upset, support others by telling them what they want to hear; the four social virtues that quietly underwrite organisational defensive routines. “The problem is not me, but you.”
Pole B. The leader’s actual behaviour, especially under threat or embarrassment, is taken as the data. The gap between espoused and in-use is the diagnosis. Under threat, Model II governs: grant legitimacy to others’ views, assume partiality of one’s own, attribute positive intent, acknowledge impact and contribution. “Aggressive and vulnerable”: strong advocacy paired with genuine inquiry into one’s own contribution.
Pole A is right for external communication: stated principles need to be stated clearly, and an organisation needs them on the record. Pole B is right for the leader’s own self-assessment, and for any conversation where defensive reasoning is already running and the next decision will repeat the last one unless the gap is named. Most leaders, by their own report later, find Pole A was running underneath their Pole B account in at least some of the territory.
In decisions: Pole A leaders judge themselves by their intentions, run after-action reviews on the what, and end disagreements with a winner; Pole B leaders judge themselves by their last three observable interactions, run after-action reviews on the what we believed coming in, and end disagreements by asking which observable example would change the other person’s view, and whether the same example would change theirs. This dichotomy is the load-bearing wall of this mirror: every other entry can be self-reported into the flattering pole, and only the gap between espoused and observed produces signal the leader didn’t already have.
Watch. Roderic Yapp, Double-loop learning: a case study from the front-line (TEDxWandsworth, 17 min): two front-line case studies, Afghanistan platoon houses and then a mental-model challenge, that name the espoused/in-use gap directly. Companion: Chris Argyris, Chris Argyris Talks About Culture and Management (4 min). The framework’s primary author in his own voice, defining theory-in-use and running the canonical “bypass the threat, cover up the bypassing, cover up the covering up” sequence.
Pole A. Strong, confident, resolute positions communicate competence; doubt is weakness.
Pole B. Lack of certainty is strength; certainty is arrogance. Think out loud; voice context and hunches; invite challenge; make 90% confidence interval predictions and audit calibration after.
Pole A is right in genuine emergencies—Chaos in Cynefin terms—where projected confidence is the team’s anchor and the cost of paralysis exceeds the cost of being wrong. Pole B is right almost everywhere else, especially in low-validity decision domains where confidence tracks coherence of story rather than quality of evidence.
In decisions: Pole A leaders make point predictions and defend them; Pole B leaders advance positions with the test that would falsify them.
Pole A. The right answer comes from listing the options, comparing them against criteria, and choosing the best. Deliberate analysis is what makes a decision defensible.
Pole B. The right answer comes from recognising the situation as one the experienced operator has seen before, generating one workable course of action, and mentally simulating it forward. Pattern-recognition is what makes a decision fast and, in the right conditions, reliable.
Pole A is right when the environment is irregular, feedback is delayed or absent, the stakes are high enough that the decision will be publicly justified, or the operator is not yet experienced enough for Pole B to be trustworthy. Pole B is right when the environment is regular enough that experience has built trustworthy patterns and feedback is rapid enough to keep those patterns calibrated: emergency response, surgical theatre, deployment rollback, anything an expert has done a thousand times. Klein and Kahneman, who spent years disagreeing about whether expert intuition was real, eventually agreed: it is real where the environment is regular and feedback is rapid, and it is unreliable everywhere else. That joint conclusion is the diagnostic for which pole fits.
In decisions: Pole A leaders introduce a checklist or formal review where stakes are high and feedback is poor; Pole B leaders trust the experienced operator’s first reasonable option and notice when the situation has slipped outside the territory the operator’s experience covers.
Pole A. Knowledge that matters is universal, codified, decomposable, teachable as formal discipline; standards, documentation, and knowledge-management systems are the organisation’s memory.
Pole B. Most operating knowledge is local, implicit, and held by practitioners: James C. Scott’s mētis, Nonaka’s tacit knowledge, the river pilot’s feel for the one harbour. Mentor-mentee transmission and shared experience are the real infrastructure.
Pole A is right where the work is genuinely procedural, the knowledge is genuinely codifiable, and personnel turnover is high enough to need it. Pole B is right wherever a work-to-rule strike would halt the operation, which is almost everywhere outside fully automated processes.
In decisions: Pole A leaders standardise practices across business units and invest in documentation; Pole B leaders ask the local team what they already do that the standard fails to capture, protect it, and invest in senior staff coaching junior staff at the work.
Pole A. Information is gathered, analysed, decided upon; the cognitive faculty is the relevant tool.
Pole B. The body knows before language does; felt sense, presence, embodied attunement are signal, not noise.
Pole A is right when the situation is tractable enough that analysis converges and the body’s signal is dominated by cognitive bias. Pole B is right when the situation is intractable, analysis has converged on a position the body finds suspect, and the suspicion is data.
In decisions: Pole A leaders convene more analysis when stuck; Pole B leaders pause, attend to body, name the discomfort, and let the next move emerge.
Watch. Eugene Gendlin, Focusing (Nada Lou archive, 12 min). Gendlin himself on felt sense: “the body has its own take on what’s going on… more subtle, more intricate… you walk by the familiar feelings… one more step.”
Authority: how does this leader read power, decision rights, and the leader’s role?
16. Leader-follower vs leader-leader (and the orders that go with each).
Pole A. Authority lives with the leader; subordinates seek permission and report status; orders are detailed, control is by supervision, deviation is failure.
Pole B. Authority lives where the information lives; subordinates state intent and act, the leader replies “Very well”; orders are intent, purpose, constraints, and antigoals, with subordinates choosing the right action as conditions change.
Pole A is right in genuine emergencies where time-to-decision is the binding constraint and command structure is what is being practised, or where the right action is precisely specifiable and the cost of unauthorised initiative is high. Pole B is right almost everywhere else, including most of what looks like an emergency, and especially where the right action depends on local conditions.
In decisions: Pole A leaders’ calendars fill with approval meetings and they write detailed task lists; Pole B leaders’ calendars fill with conversations that begin “I intend to…” and they write a paragraph of purpose and a list of what must not happen.
Pole A. Information flows up, decisions flow down; the org chart is the operating system.
Pole B. Information flows where it is needed; decisions are made where the information is; the org chart is one of several views.
Pole A is right where regulatory accountability or genuine command structure requires it: defence, regulated finance, certain public-sector contexts. Pole B is right almost anywhere else, including in regulated industries where the structure has been kept for inertial reasons.
In decisions: Pole A leaders re-org by redrawing reporting lines; Pole B leaders re-org by changing meeting cadence, information flow, and decision rights. Steve Jobs named the Pole B move plainly at D8 2010: Apple is “organised like a startup… run by ideas, not hierarchy.”
18. Theory X vs Theory Y. See the worked example above. Pole A is right in two distinct contexts: genuinely high-stakes regulated work with catastrophic-bad-actor risk (pharmaceutical batch release, nuclear plant control rooms, anti-money-laundering compliance, aviation cockpit procedure), and routine low-skill, low-discretion work where the task itself doesn’t generate intrinsic motivation and the structure is doing the load (high-turnover transactional roles, certain warehouse or contact-centre tasks where the work is the work and capability development sits outside it). Pole B is right in most knowledge work, most cross-functional work, and most situations where intrinsic motivation is doing more of the load than the control structure can see.
Pole A. The leader’s authority and competence drive the team; subordinates execute; the leader takes personal credit for success and personal responsibility for failure; the organisation orbits the leader’s competence.
Pole B. The leader retains final authority but creates highly participative teams; builds capacity that outlives them; the test of leadership is what happens after they leave.
Pole A is right in genuine crisis where personal visibility holds the team together, and is over by the time the crisis ends. Pole B is right almost everywhere else.
In decisions: Pole A leaders close debates by deciding, accept extension after extension because the organisation “needs them”; Pole B leaders close debates by surfacing the disagreement and inviting the team to decide, and measure success by their successor’s competence.
Pole A. Organisation as machine; meritocracy, MBO, KPIs, R&D, balanced scorecards; effectiveness as the yardstick.
Pole B. Organisation as living system; self-management, wholeness, evolutionary purpose; sense-and-respond rather than predict-and-control.
Pole A is right in tractable, stable industries with well-understood competitive dynamics: regulated utilities, mature consumer staples, infrastructure with predictable demand. Pole B is right in industries undergoing genuine technology-driven disruption where the planning horizons no longer track the disruption rate, which is most software-permeated industries right now.
In decisions: Pole A leaders run quarterly cascades of OKRs; Pole B leaders run quarterly listening practices to “what the organisation wants to become.”
Pole A. The leader’s job is to know the answer and communicate it clearly. Effective leadership is articulate direction; effective change is handing the team the answer, the plan, the playbook.
Pole B. The leader’s job is to surface the information needed for good decisions. Ask questions to which you do not know the answer; let the team find the answer; develop capability, not just compliance.
Pole A is right when the leader genuinely knows the answer, the team genuinely does not, and either the cost of getting the answer wrong is high or the team genuinely lacks the prerequisite skill and the cost of failure is high. Pole B is right whenever the leader is acting from confident position on a question the team is closer to than the leader is, or whenever the leader is solving problems faster than the team can grow into.
In decisions: Pole A leaders open meetings with a position statement and shorten meetings by deciding faster; Pole B leaders open with “What’s on your mind?” and lengthen meetings by asking the five Toyota Kata questions, accepting that the team’s answer may be better than theirs.
Pole A. The leader is the most senior technical authority; problems are escalated up to where the expertise is.
Pole B. The leader’s role is to make others’ decisions possible; expertise lives at the work.
Pole A is right when the leader is genuinely the deepest expert and the call is technical: the founding CTO making the last call on a core architecture decision, the chief surgeon overruling on the table, the lead investigator on a specific case. Pole B is right when the leader’s expertise is a layer or two removed from where the work is actually being done, which is almost every senior leadership context outside narrow technical-founder situations.
In decisions: Pole A leaders are the technical decider on contested calls; Pole B leaders convene the people who know and facilitate the call they make.
Causation: how does this leader read the relationship between events and their causes?
23. Deviation vs conditions.
Pole A. When work goes wrong, the explanation is a deviation: from procedure, from competence, from intent. The fix clarifies the standard or addresses the deviator. “Human error” is a sufficient diagnosis.
Pole B. Real work continually adjusts to underspecified conditions; the gap between work-as-imagined and work-as-done is where both safety and risk live. Failure is the unexpected combination of normal variability; “human error” is a symptom of that gap, not a diagnosis.
Pole A is right in tractable Complicated/Ordered failures: a broken bearing, a software null-pointer exception, a structural calculation error, an individual whose behaviour was genuinely reckless. Pole B is right in any incident where the failure was the unexpected combination of normal variability, which is most of what goes wrong in Complex/Unordered work.
In decisions: Pole A leaders write performance-management documents, root-cause reports, and policy clarifications; Pole B leaders go to the gemba (the Japanese term for “the actual place”, where the work happens), ask “what surprised you? where did you have to improvise?”, and change the system people are operating within.
24. Safety as absence vs safety as presence. This is a corollary of entry 23, framed at the level of metrics and investment rather than incident analysis.
Pole A. Safety is the absence of events; the fewer accidents, incidents, and near-misses, the safer the system. Safety investment is insurance against accidents.
Pole B. Safety is the presence of defences, the ability to succeed under varying conditions. Safety investment is investment in productivity.
Pole A is right in tractable systems with stable specifications: the lost-time injury rate on a mature manufacturing line, the field-failure count for a long-standing physical product, the audit-finding count on a stable regulated process. Pole B is right in intractable socio-technical systems where work-as-done routinely adjusts to conditions work-as-imagined did not anticipate: a software platform under continuous deployment, a clinical service operating with staffing variability, anything where the absence-of-events metric stops moving while everyone reports getting away with more than they used to.
In decisions: Pole A leaders count injuries; Pole B leaders ask “how are we able to do this work well 999,999 times out of a million, and what do we need to preserve about that?”
See also. The sources under entry 23, Dekker’s Safety Differently and Allspaw’s How Your Systems Keep Running, both teach this dichotomy at the metrics-and-investment level too.
Pole A. Beginning-middle-end arcs with clear heroes, villains, and lessons.
Pole B. Several stories simultaneously about the same event; the discomfort of multiple frames as the data, not as a sign that the analysis is incomplete.
Pole A is right when the audience needs to act and the action is genuinely unambiguous: the all-hands message after a successful product launch, the board narrative on a clean quarter, the customer communication after a single root-cause incident with a documented fix. Pole B is right whenever a confident single-cause story is being constructed under conditions the data does not support: most strategic narratives, most multi-quarter performance explanations, most post-incident reviews where “human error” was offered as the answer in the first 24 hours.
In decisions: Pole A leaders narrate quarterly results as a story of decisions made; Pole B leaders narrate them as a story of conditions encountered, with several versions on offer.
Watch. Garvey-Berger, Mindtraps — Simple Stories (5 min). Companion: Tyler Cowen, Be suspicious of (simple) stories (TEDxMidAtlantic, 16 min): names the narrative fallacy directly and argues for “the mess” (multi-causal reality) as the alternative to single narrative. “Every time you tell yourself a good-vs-evil story, you’re lowering your IQ by 10 points.”
Self: how does this leader read their own identity, capacity, and developmental edge?
Every other dichotomy is mediated by the leader’s relationship to themselves. The developmental literature drawn on here (Kegan, Lahey, Garvey-Berger, Joiner) has decent measurement reliability for trained scorers, modest construct validity, and contested predictive validity. The dichotomies below name observable behavioural patterns rather than relying on staging interpretations to do the diagnostic work.
26. Socialised vs self-authoring vs self-transforming.
Pole A. Identity is given by the surround: role, profession, expectations.
Pole B. Identity is a self-chosen stance; the leader can hold their own values against pressure from valued others, and the role they hold is one expression of who they are, losable without the self being lost.
Pole C. Identity is held lightly; the leader values their own frame and watches for the data that would refute it.
Pole A is the right response in some contexts: early-career apprenticeship, certain professional-formation contexts, communities of practice where the work depends on shared identity. Pole B is what most senior leadership roles structurally require. Pole C is rare and not necessary for most roles, but is the right fit in the most VUCA situations where holding any paradigm too long is itself the risk: what Meadows called transcending paradigms, the highest leverage point in her hierarchy.
In decisions: Pole A leaders cannot disappoint valued authorities on behalf of their own judgment, and cannot imagine succession because the role is doing the identity-work; Pole B leaders can disappoint valued others, and treat succession as an explicit part of how they hold the role; Pole C leaders can disappoint themselves on behalf of evidence that contradicts their judgment. A three-pole entry; the gap between A and B is usually more salient than between B and C.
Identity-given-by-surround is the structural failure mode AI exposes most cleanly. Pole A leaders cannot distinguish good from plausible: the LLM’s confident prose lands in the same register as the voices that shaped their expectations. They accept what the model tells them to do not because they’re naïve but because they have no felt sense of what fitting would mean; the model’s output is just another voice from the surround. Pole B leaders feel the dissonance when a confident output doesn’t fit the situation, and slow down. Pole C leaders treat every model output as a hypothesis to be tested against evidence the model couldn’t have seen. The leader who cannot disappoint a confident voice cannot govern an AI either.
27. Subject to assumptions vs object of assumptions.
Pole A. The leader’s frame is what they look through; it is invisible to them.
Pole B. The leader can look at their own frame; they can hold it up to inspection.
Pole A is the structural default; every leader is functionally subject to most of their assumptions most of the time, and a working organisation depends on the leader operating from a frame rather than constantly inspecting one. Pole B is right where the situation is already telling the leader their frame is not fitting (results recurring, conversations repeating, the same diagnostic appearing in different vocabulary each quarter). The diagnostic move is the noticing in those moments, not the constant inspection.
In decisions: Pole A leaders are surprised by their own pattern when it surfaces; Pole B leaders can name the anxiety that drives the behaviours they keep noticing in themselves.
Watch. Kegan, An Evening with Robert Kegan and Immunity to Change (Boston College OD Network, 14 min). Opens with the 14-frogs-on-a-log puzzle (gap between deciding and doing) and names the immunity-to-change framework. Walks through identifying big assumptions, then lands on “being able to look at the whole thing”: the subject-object move on one’s own frame.
Pole A. The decision must yield this result; if it does not, the decision was wrong.
Pole B. The decision was a test of a hypothesis; the outcome is data; the question is what to learn from it.
Pole A is right when the decision was a genuine commitment to a specifiable outcome and the accountability frame requires defending the commitment. Pole B is right when the decision was a hypothesis dressed up as a commitment, or when defending the commitment is doing more work for the leader’s identity than for the outcome.
In decisions: Pole A leaders defend the decision after it fails; Pole B leaders refine the model after it fails.
Watch. Astro Teller, The unexpected benefit of celebrating failure (TED, 16 min). Kill-and-pivot stories from Alphabet’s X: vertical farming, cargo blimp, self-driving cars (the supervised model failed, the team switched to fully autonomous). “We bonused every single person on teams that ended their projects.” Closes on “enthusiastic skepticism is optimism’s perfect partner.”
The fifteen-minute version, named at the top, is enough on its own: look back at the six axes, notice where you sit on the left, and ask where those assumptions came from and whether they still fit. If the noticing surfaces something that wants more than a glance, the two-week version below is built for that. Four movements, four calendar slots, about two weeks end-to-end.
Movement 1, the reading pass. About 1 hour, today. Read the essay through once and flag the three or four entries that strike you on first pass as most provocative for your current situation. Then sample the embedded videos and essays under the Watch/read lines for those entries; most are 5-15 minutes, a few are longer. Notice which of your initial flags shift, drop, or sharpen once you’ve watched the videos or read the excerpts for those entries. The settled list is what you take into Movement 2. The point of this step is to keep this mirror itself from doing too much of the work; the sources are doing the teaching, and your first read should reflect what you actually find when you let them.
Movement 2, the dichotomy walk. 20 minutes alone, this week. Take your settled three or four entries into solo work. For each, ask three questions: which pole am I claiming?, which pole would my last three decisions in this territory actually show?, which pole does the situation I am facing actually require? The third is the Cynefin discipline: the situation determines the right pole, not the leader’s preference. The first two are the espoused/theory-in-use discipline (entry 11, the load-bearing wall). The interesting territory is wherever those three answers diverge. Write down what you find; you will need it for Movement 3.
Movement 3, the witness round. Send by Friday; one-week response window. Send the same three or four entries to three or four people who have watched you make recent decisions in the territory. Ask them which pole they observed in your last three decisions. One sentence each is enough. Compare to your own answer from Movement 2. The gap, when it appears, is the diagnostic.
Movement 4, the conversation. 45 minutes with a coach, the week after. Take the gap to a coach and work it: make visible the operating pattern that is already happening, and make the next decision a choice rather than a reflex. The achievable end state is one sentence you can say out loud: “On entry N, my espoused pole is X, my observed pole is Y, and the situation usually calls for Z.” If you do not yet have a coach who works at this depth, finding one is the work; this mirror is built for that conversation, and the conversation is what makes this mirror do anything other than sit on the page.
A staff-runnable variant. For a leadership team rather than an individual: same four movements, no coach. Movements 1 and 2 stay solo. Movement 3 is the team being each other’s witnesses on a single shared entry (entry 1 is a strong first choice). Movement 4 becomes a 30-minute conversation about the gaps that surfaced: what each person’s espoused pole was, what their last three decisions in this territory actually showed, and what the situation the team is facing actually requires. Run on one entry per fortnight. The team’s situation, not the facilitator’s, determines what each conversation produces.
What this mirror does, and what it doesn’t
This mirror makes the operating paradigm observable. It does not, on its own, change it. The change-work is the conversation that follows. The mirror is the first step of that work, not a substitute for it.
Three things you might do next.
If one entry hit hard: send it to one peer who has watched you make recent decisions in that territory, and ask them which pole they observed. One sentence is enough. That single exchange will tell you more than another read of the article will.
If you want to talk: flick me a note on LinkedIn, or visit hi.chrisgagne.com. If you do not yet have a consultant and/or coach who works at this depth, finding one is the work.
Some book links are Amazon affiliate links. If you buy through them I earn a small commission at no cost to you. Primary-source citations (papers, reports, articles) link to the original or to free archives.
This essay grew into a book.Come Prepared to Die is now being written and serialised chapter by chapter: on LinkedIn in the Come Prepared to Die newsletter and here on the blog under Come Prepared to Die. What follows is the original essay the book grew from. Subscribe on either to follow along.
TL;DR: The paradigms most organisations run on were survivable when machines couldn’t replace people. AI changes that. Almost every transformation of the past several decades has been structurally capped at what the leader’s paradigms permit, and the paradigms that got leaders here—Theory X, utilisation, plan-tyranny, and others—are now lethal rather than merely sub-optimal. The work that’s needed is neither more consulting nor more AI on top of the existing paradigm. It is coaching of a particular depth—rooted in presence and refusal to shortcut—held inside a support community, because leaders at the level of paradigm-death are moving through a rite of passage that cannot be done alone. This essay is the structural argument for why, and a practitioner stance for what comes next.
NB: I am standing on the shoulders of many giants and this essay is a small increment upon them. Most of this work is not original in isolation, but I believe the synthesis is. By the end you’ll know what to do—though not necessarily how to do it—and who to turn to for help. This work cannot be done alone from within the existing paradigms and it takes more than a typical consultant’s help. Who’s in?
1. The slow dawning
I did not have a single clarifying event. Drift, in Dekker’s sense. The cultural shorthand is the boiling frog: wrong about frogs, folk-right about this.
The clearest event landed a few weeks ago. I gave Claude a piece of work that would have taken me several months: regenerating around 350 reference documents I maintain, using a nine-pass ingestion protocol that audits each output against the original source. A day and a half later it was done.
Christophe Louvion had said it to me in his usual one-line form a bit earlier: the LLM does not have to be totally correct, it just has to work at least as well as the average human. I heard the sentence then and not yet felt it until now.
I have spent much of my life relying on intellect and ways of seeing the world for protection in a volatile, uncertain, complex, and ambiguous world. Now, faced with my vulnerability—our vulnerability—I am left with a choice: keep trying the same approach and die of exhaustion, or give up and attend to what is here. I wrote about that side of the passage in Just This, also on LinkedIn, earlier this month.
Like most of us, I have been here before in smaller versions. Serious adverse childhood experiences. A father, diagnosed with cancer, whom I could not visit for most of his last two years because of New Zealand’s COVID-era travel restrictions. Each of those was a death and rebirth in its own way.
My life at the moment is a crucible of a chrysalis. This is not a phase I can exit by reading the right things or seeing a therapist for a year or two. It requires a change in identity.
My Alētheia coaching teacher, Steve March, named the larger register on a coaching call this month. Sometimes we get lazy, he said, because we think tomorrow will look more like yesterday than something different. Six months from now, a year from now, we may be living in a different sort of world.
The version of the recoil I have been living, somewhere between Steve’s call and Moore’s insight, is the experience of waking on an ordinary morning to discover that the world has fundamentally shifted and appears the same. The not-knowing is the anxious thing. There is nobody I can turn to who has been in this context before, because none of us have. A crisis is a clarity call, a moment to wake up and change our paradigms. It’s time for us frogs to jump out of the boiling pot. The water has been getting hotter for decades—more volatile, uncertain, complex, ambiguous—and AI is the rolling boil.
Some version of that context is probably what brought you here, whether you are in crisis or have the foresight to see one coming. The work of transforming organisations and the work of transforming lives are not different questions.
2. The 10% ceiling
Almost every Agile transformation of the past 25 years has failed relative to what its framers intended. I have a 20-plus-year career in this work. I led transformations I am proud of, and I participated in or led others I called meaningful change at the time and now read more soberly. None have fully met my aspirations. In some, it might be that the only real success was inoculating the rank-and-file leadership I could access against cargo-cult behaviours, even if I could not shift the broader organisation.
Craig Larman noted in his Certified LeSS Practitioner (CLP) course that where transformation does (rarely) succeed, it lasts at most seven years before a change in executive leadership rewinds it. If a transformation looks like a longer success than that, the leaders and their consultants are likely lying to themselves about what has actually changed.
Most consultants work at what Peter Block calls the content level: technical recommendations, frameworks, deliverables the manager can absorb without changing how the manager sees themselves or their organisation. There is also an affective level: trust, power, identity, the relationship between consultant and client and between leader and led. The content level is where consultants are trained. The affective level is where paradigm-change has to happen. Most consultants will not work at the affective level because it threatens the consultant’s position as much as the manager’s. Consultant and manager collude on keeping it that way. Block names the collusion: too often we collude with the client in pretending that organizations are not political but solely rational.
Ronald Heifetz, later with Linsky and Grashow, names the same gap structurally. Technical problems have known solutions inside current authority and expertise; adaptive challenges require changes in values, beliefs, roles, and relationships, and the most common failure of leadership is applying technical fixes when the work is actually adaptive. The 10% ceiling is what consulting does when it stays inside technical mode and the work is adaptive.
Robert Kegan and Lisa Lahey give the gap its mechanism. Change rarely fails for lack of sincerity or willpower. It fails because we mean both things at once: a sincere commitment to the goal, and a hidden countervailing commitment to self-protection, sustained by big assumptions held as fact. The leader’s paradigm is exactly this kind of immune system, and it is not weak or sloppy; it is a brilliantly designed protective architecture, built at a time that may no longer fit.
Jerry Weinberg turned the boundary into practical advice: Never promise more than 10% improvement, because most people can successfully absorb 10% into their psychological category of “no problem.” Anything more, however, would be embarrassing if the consultant succeeded. The figure has held for forty years, not because it is empirically calibrated, but because it is the boundary at which the consultant does not threaten the manager’s paradigm. Block, Heifetz, and Kegan-Lahey describe the same boundary from three angles: Block from the consultant’s side of the collusion, Heifetz from the kind of work the boundary excludes, Kegan-Lahey from the immune mechanism that holds it in place. Weinberg’s 10% is where the four lines meet.
Alexey Krivitsky, Craig Larman, and Roland Flemm, in their excellent 10X Org, return to Weinberg’s rule and argue we cannot afford the ceiling any longer. I think they are right.
Eli Goldratt, in The Choice, made the cognate cut 27 years after Block. He named four obstacles that keep people from clear thinking. People believe reality is complex; conflicts are a given; other people are to blame; they already know. These are paradigm-level obstacles, not technical barriers better methods could remove. They are the fundamental orientations a person takes toward reality itself. Goldratt’s working assumption, what he calls Inherent Simplicity, is that reality, any part of reality, is governed by very few elements, and that any existing conflict can be eliminated, if the person looking is willing to set aside the four obstacles. Almost no leader is. The four obstacles are how the leader’s paradigm protects itself.
The aviation version has been visible for decades. Gladwell’s popular reading of Korean Air Flight 801 placed the cause inside the Korean language’s grammaticalised speech levels. The harder reading that the NTSB stopped short of is that linguistic deference made the gradient more visible without causing it. The cockpit voice recorder shows the first officer and flight engineer questioned the captain about the glideslope and registered alarm at the GPWS callouts, but never escalated to a clear directive. They were correct on the merits, and tentative in the framing. The first officer finally called for a missed approach about seven seconds before the aircraft hit Nimitz Hill. The crash killed 228 people; one further survivor died later.
The deep variable was the steepness of the authority gradient, not the grammar. Aviation has spent four-plus decades engineering the gradient down through Crew Resource Management. Knowledge-worker organisations mostly haven’t, which is why the same dynamic still shows up in retrospectives, failed product launches, unspoken concerns about a CEO’s pet project, and board rooms where everyone knew the strategy was wrong and nobody said it directly enough. Amy Edmondson’s nearly three decades of psychological-safety research document the boardroom version of the gradient.
Block’s ceiling, Heifetz’s adaptive challenge, Kegan-Lahey’s immune system, and the cockpit’s authority gradient are all describing the same ceiling from four different angles. The cockpit version has been investigated, regulated, and partly fixed; lives so visibly at stake make complacency hard. The boardroom version has not meaningfully changed. The asymmetry is the explanation: the diffuse, statistical harm of a misfit paradigm at scale is not as visceral as a plane crash, even when its aggregate cost is larger. Visible mortality forced aviation to engineer the cockpit’s gradient down. Mortality, in indirect and slower form, is now visible in the boardroom too. AI is what is making it as visible as a plane with all of its engines on fire, the early smouldering not recognised soon enough because the plane still flew.
Thus, almost every transformation of the past several decades has been structurally capped at the boundary the leader’s paradigm permits. Not because of bad execution. Not because Agile or Lean or DevOps were wrong. Because the leader’s paradigm could not be approached, and the consultants in the room either could not—or would not—approach it. This is neither the leader’s nor the consultant’s fault, but it is our shared responsibility.
Almost every transformation of the past several decades has been structurally capped at the boundary the leader’s paradigm permits.
3. Some misfit paradigms
What follows is a partial catalogue for complex knowledge work: creative, ambiguous, volatile. A help-desk operation can run on Theory X with utilisation targets and not be in crisis, even if it is a pale shadow of what is possible. The paradigms below break down where ambiguity, learning, and human judgement are load-bearing. These were the right paradigms when the water was cooler, when the work was more predictable, the customers more uniform, the timelines longer. The water has been heating for fifty years; AI is the moment it boils. Mary Poppendieck named several, plan-tyranny most pointedly. Each is still operative inside most organisations I work with. To declare them as immutable or a commercial reality is practically a thought-terminating cliché.
AI is the moment our existing paradigms have become terminal, not merely sub-optimal.
Theory X. Douglas McGregor’s distinction between Theory X (workers cannot be trusted; require coercion and control) and Theory Y (work can be as natural as play; people seek responsibility and meaning) is 66 years old. Most organisations still operate on Theory X assumptions even when leadership quotes Theory Y. Bas Vodde taught me an exercise in his CLP course that catches this in flight. Draw a horizontal line. Theory X to the left, Theory Y to the right. Sticky-notes above the line for people’s instincts and behaviours, below for the structure and processes they are operating inside. If you ran this exercise, what would you see in your org? In our course, we all saw lip-service Theory Y, structural Theory X. I’ve not seen an exception yet.
High utilisation, compounded with hyper-specialisation. Donald Reinertsen’s The Principles of Product Development Flow gathers half a century of queueing theory to make a single point. Keep-everyone-busy makes everything slower. High utilisation produces queues, queues produce delay, delay produces variability, variability produces yet more queueing. The 10x organisation literature and Org Topologies catalogue the second half of the same dysfunction. Treating people as interchangeable resources—meat widgets, in Larman’s words—slotted into functional silos, produces dependencies, hand-offs, and waiting times. The coordination cost scales worse than the labour cost it was meant to optimise. Hyper-specialisation creates more queues for value to flow through; utilisation makes each queue at least an order of magnitude longer. The two compound.
The tyranny of “the plan.” Mary Poppendieck named this one and I have not found a better name for it. PERT charts were invented for the Polaris submarine programme in the late 1950s, not because PERT worked, but because Rear Admiral Raborn needed something he could show Congress to keep an eight-year programme funded across multiple election cycles. The planning grammar of the modern project-management profession was built on top. Harvey Sapolsky’s 1972 history of the programme confirms the bureaucratic reading: the technical officers bypassed PERT, the contractors considered it worthless, and the charts functioned as a façade for the funders. The façade became the doctrine, and a generation of project-management training, PMI’s included, was built on top of it.
The mechanism the doctrine encodes is straightforward. You decompose the work, sum the estimates, and, as Poppendieck puts it: bingo, there’s your schedule. You manage to it. When reality predictably diverges, the planning paradigm has only one reading ready: someone failed to try hard enough. The alternative reading, that the plan was merely a hypothesis and the divergence is useful information, is not structurally available. Variance-from-plan is the signal; conformance is the definition of success; and the harder people try to execute a wrong plan, the less the organisation learns about why it is wrong. This is Theory X dressed up as project management.
Other paradigms warrant the same examination, this list is not exhaustive. The smart-technocrat reflex. Konosuke Matsushita warned Western managers in 1985 that “the intelligence of a few technocrats has become totally inadequate.” AI now offers an industrialised version of the same paradigm. Leader-follower hierarchies, the assumption-of-conflict between management and workers, the myth of the heroic individual, bonus and incentive cultures that crowd out intrinsic motivation. The lineage runs from Senge through Argyris, Schein, Kegan, and Heifetz; each has been visible for decades, written about in books still in print, named long before AI, and is now turned terminal by it.
Donella Meadows placed the power to transcend paradigms at the top of her catalogue of leverage points, above paradigms themselves, above goals, above information flows. She wrote: there are no cheap tickets to mastery. You have to work hard at it, whether that means rigorously analyzing a system or rigorously casting off your own paradigms and throwing yourself into the humility of Not Knowing.
These paradigms aren’t wrong everywhere. Most have contexts where they fit. If a leader ran one for thirty years it isn’t because they were stupid, it’s because the context rewarded it. Alexey Krivitsky, Craig Larman, and Roland Flemm cemented the historical case for me via the canonical example, foregrounded in 10X Org: Adam Smith’s pin factory (1776), where ten workers using division of labour produced 48,000 pins a day against a solo baseline of one to twenty per worker, a roughly 480-fold gain. The reason it worked is that a pin factory operates in a low-variability environment. Every pin is identical to the last, so hyper-specialisation and high utilisation compound into vast productivity rather than into queues. Mary Poppendieck made the same move from the other direction: in manufacturing, variation is the enemy; in service, variation is required because customer expectations differ .The paradigms aren’t wrong; they were lifted across a domain boundary they did not fit. Hammers work for nails—not screws—but they are still useful tools.
When the underlying paradigm remains the same, performance drifts back toward the old baseline. None of the structural redesign survives the leader’s unreconsidered paradigm. None of which is to fault leaders for having those paradigms; the paradigms got most of them their careers. They worked in cooler water.
4. Why the ceiling holds: paradigm-change is identity-death
The leader can see the catalogue intellectually. So why doesn’t the leader simply move? Ram Dass put the answer plainly in Be Here Now, 55 years ago. In the process you must die. This is not a physical death, but it is the end of the world as we know it.
Changing one’s paradigm is deeply vulnerable because it is a shift at the identity level for most of us. If I am not the leader with the answer, am I powerless? The paradigm is the conventional self. To change it, even partially, is something close to ego-death. Serious contemplative practitioners can take years to intellectually understand they are not their thoughts, and even longer to embody it. Asking a senior leader to do the same on a fixed “transformation project” timeline in the midst of their daily work is not a practical ask.
The self that built the career, the self the paycheque pays. All of it is held together by the collection of paradigms and thoughts we call ourselves. Jennifer Garvey-Berger names the same trap directly: shackled to who you are now, you can’t reach for who you’ll be next. Ram Dass said the same in Be Here Now: it’s only when caterpillarness is done that one starts to be a butterfly. And that again is part of this paradox: you cannot rip away caterpillarness. The whole trip occurs in an unfolding process under which you have no control. Chris Argyris named the same thing from the cognitive side in the 1970s: single-loop learning changes behaviour while leaving the underlying assumptions intact, and the assumptions don’t move because they are the self.
W. Edwards Deming named the pace bluntly: this work is not fast and it cannot be bought, only facilitated. The same lineage carries another line: come yourself, or send no one. It’s associated, possibly apocryphally, with Deming’s reply to the Nashua Corporation CEO who asked to learn his methods through a delegate. The structural point holds either way. Most leaders will send someone, often their VPs. The transformation never quite arrives, because the ultimate leader did not.
The CEO is running paradigms that are not fit for what AI is making general. There’s no malice nor stupidity here; this is structural. The conviction is communal: most of their peers think identically. Who goes out on that ledge alone? Beatrice Bruteau named this in The Psychic Grid: the world we know is one we collectively project, and few of us can step outside the projection alone. The leader got to the corner office doing what they are doing, and by all comparisons they are doing fine, because the floor is just that low.
The actuarial reality. It is not that self-driving cars do not make mistakes. It is that they make fewer than humans on average, even when the mistakes they do make look horrifyingly dumb by human standards: two different brands of driverless cars drove into clearly marked wet cement. No human driver paying attention does this. By January 2025, Waymo had logged 56.7 million rider-only miles, with an 85% reduction in suspected serious-injury crashes and a 96% reduction in injury-involving intersection crashes versus aligned human benchmarks (Kusano et al., Traffic Injury Prevention, 2025). The actuaries know what they are looking at. LLMs are at the same threshold for cognitive work. “Computer” was a human job title for over 300 years before machines inherited the word. Actuaries see this in driverless cars. Christophe saw it with AI and many others are catching up.
The augmentation cost. Lisanne Bainbridge named this in 1983: automation takes the easy parts of the operator’s task and leaves the hard parts unsupported, then asks the operator to step in when the system fails. The operators most degraded by automation are the ones expected to catch what it misses. A healthy person chooses an electric scooter for short distances because it’s faster with less effort. They lose the muscles and the cardiovascular fitness, then the proprioceptive intelligence, then ultimately the sense of their own embodied capacity. Twenty years on, the scooter is no longer a choice; the body cannot make the trip. We are doing this with cognition, deliberately, at scale. Lee and colleagues at Microsoft Research and Carnegie Mellon, in an early-2025 survey of 319 knowledge workers, found confidence in AI tracked inversely with critical thinking. An MIT Media Lab team running EEG sessions on LLM-assisted essay writers (Your Brain on ChatGPT, 2025, n=54) recorded reduced neural connectivity, poorer recall of the writers’ own prose, and homogenised output among heavy users. A 666-person study by Gerlich in Societies the same year found AI-use frequency inversely correlated with critical-thinking scores, sharpest among the youngest. These studies measure averages. For a smaller group of users who treat AI as a thinking partner with critical engagement—who push back, ask for sources, and refuse first answers—the opposite may hold. The displacement of cognition likely tracks the manner of use, not the tool.
The paradigm picture
Universalisability. The claim most AI commentary ducks. Imagine a single firm that has built extraordinary value by replacing knowledge workers with agents. Stable, profitable, defensible, locally optimal. The fabled 1-person, $1B company. Now universalise: what happens if every firm does the same? The aggregate is that the workers have no wages to acquire the extraordinary value the firms are producing. The firms have built a position that destroys the conditions for their own continuation.
Kant’s categorical imperative—act only according to that maxim whereby you can at the same time will that it should become a universal law—is not only moral fashion. It is a stress-test for how a position holds when universal. The corporate AI position fails the test, and most leaders deploying it know it fails. The game theory says they must attempt it anyway. Each firm is acting rationally; the aggregate is collectively suicidal. Hemenway Falk and Tsoukalas formalised this as a fallacy-of-composition with a demand externality in The AI Layoff Trap (arXiv, March 2026). The same pattern shows up at the supplier level: a handful of frontier-model providers concentrating strategic power across the entire knowledge-work economy, with each firm rationally choosing dependency and the aggregate producing a single chokepoint no firm intended to create. It is the structural shape of every commons collapse, and it is happening now in the most consequential industry in the world. The leaders who weather this best are the ones who can hold, simultaneously, that they must compete and that the competition has a destination none of them want to arrive at.
Each firm is acting rationally; the aggregate is collectively suicidal.
The disintegrated moral structure. When IBM was the dominant computing firm, it held a moral floor: the line attributed to a 1979 IBM training slide, A computer can never be held accountable, therefore a computer must never make a management decision. That floor has been honoured in most large firms for most of my professional life. The floor has disintegrated. Daron Acemoglu’s The Simple Macroeconomics of AI (2024) gives the disintegration a quantified line: the bad-task externalities of AI (manipulation, deepfake, security-arms-race revenue) could appear to raise GDP by about 2% while reducing welfare by roughly 0.72% in consumption-equivalent terms. Small in absolute terms; large as a direction. The floor held when machines could not replace people. Now they can, partially.
LLMs drown out the source on leader paradigms. Two of Craig Larman’s laws of organisational behaviour apply directly to what AI is now industrialising. Corollary 2 is the broad claim: any change initiative will be reduced to redefining or overloading the new terminology to mean basically the same as status quo. Corollary 4 is the specific mechanism: displaced managers and single-specialists become “coaches/trainers” for the change, frequently reinforcing the terminology hijacking in their writing. My own corollary, built on Larman’s laws and worked through over a few years: the original insights of any change movement, Agile, Lean, DevOps, were always a small minority of the corpus, and the training data is dominated by that consultant-derivative watering-down. When an LLM writes on these topics, it regresses to the mode, not the truth. If almost everyone writing about Scrum Masters is describing project managers, the model will define a Scrum Master as a project manager, no matter what Sutherland and Schwaber intended. An LLM has statistics, not discernment. Chris Argyris and Diana McLain Smith named the human analogue half a century ago: defensive routines persist because we support others by telling them what they want to hear. Put those together and the leader-paradigm-specific version follows.
AI does not solve the leader-context-paradigm mismatch. It sells it back to you. The model is not just regressing to consultant content in general; it is codifying the dominant worldview and feeding it back to this particular leader with affirmation, in the precise places where critical thinking is most needed. Most leaders deploying agents at scale are buying a tool that codifies the paradigm they need to leave, with a fawning persona that promotes continued usage rather than critical thinking.
I’ve tested this directly. Asking Claude “what is Agile?” produces consultant-derivative ceremonies-and-frameworks. Even when I asked for research, Claude told me the question didn’t require it. So I built a different retrieval architecture (soon to be open-sourced under MIT) that pre-projects source references onto specific task domains at ingestion, under a nine-pass protocol with a source-only audit gating every output. Asking the same model the same question—this time grounded in that library—produced something else entirely: Agile as a learning discipline, grounded in the Poppendiecks, anchored to Theory Y via Larman, with the AI-era constraint migration named explicitly. Totally different paradigms.
AI does not solve the leader-context-paradigm mismatch. It sells it back to you.
The mētis ceiling. James Scott’s distinction between techne (codifiable, transferable, formalised knowledge) and mētis (situated, practical, implicit, partly tacit knowledge) marks another structural limit. Every functional organisation runs on more mētis than its leaders know. The LLM cannot acquire mētis because mētis is not in the corpus. Tooling-side discussions of this limit have run since 2023. The paradigm-level framing is different: mētis is a structural ceiling on how much paradigm-misfit AI can mask before the system starts to fail. Current discussion of context and harness engineering is correct as far as it goes; the harder question is whether any harness can be sufficient if it cannot contain the human’s mētis. I do not think it can. The AI optimist’s answer is that mētis will eventually be encoded; the optimist’s honest answer is that we do not know how, and the encoding might not survive being made explicit. Following Eugene Gendlin, I am less swayed: you can say and think a lot for years, knowing all the while that you aren’t touching what is implied. What does a strawberry taste like? It can manipulate text about the felt sense; it cannot register the implicit in its own body, because it has no body.
Why the productivity-overwhelm objection doesn’t escape the ceiling. The strongest objection runs roughly like this: AI productivity gains will be so large, with some forecasts running to 100x, that the firms displacing workers can absorb the political backlash with the wealth they capture. Harry Holzer documents that demand effects have historically absorbed automation-driven productivity gains across the Luddites, 1950s computerisation, and China outsourcing; forecasters consistently underweight the income-and-spending response. (Daniel Vacanti shows arrival rate routinely exceeds completion rate in real organisations; the gap is often two to four times, partly because adding capacity induces its own demand.) Put together: a 2-10x AI productivity gain just clears the present backlog. It does not displace the workforce that was generating, prioritising, and resolving the work in the first place. The displacement story confuses output volume with employment ceiling. Any one of the paradigm shifts catalogued above (abandoning utilisation, breaking hyper-specialisation, treating plans as hypotheses) could produce 2-10x on its own, and has, in the organisations that have actually done it. The AI-only equilibrium does not hold against the same queueing economics that broke the utilisation paradigm. The ceiling moves; it does not vanish. (The honest caveat: if knowledge-worker labour displaces faster than the political economy can absorb, we will be in a crisis larger than any single firm can navigate, and none of this matters at the system level. That is its own essay.)
6. The work that is needed
Steve March offered a frame on a recent coaching call: what if none of it was wasted? Not the years inside the misfit paradigms. Not the transformations that didn’t take. Not the decades of running a playbook that no longer fits. The path is fits and starts, mistakes, steps backward. The life lived exactly as it has been lived up to this moment is the doorway. There is no other.
The catch is the slow work. The opportunity, paradoxically, is also the slow work: a once-in-a-career window for the leaders who can refuse the false quick win and let caterpillarness happen.
On the practitioner side, the diagnostic craft survives the passage: named-pattern recognition, harness work that runs without me in the room, frameworks that let mid-tier authors do leadership-grade systemic learning. What does not transfer to the harness is the rest: two solid days of in-person workshop with a team produces something in their bodies and in the room between them that no agent can reproduce in years of asynchronous work. The agent isn’t bad at it. The thing transmitted in the room is not made of words. The agent works in words. The room works in something more.
An atlas has value. An atlas does not replace a year spent living and travelling in a foreign country.
The figure
What is needed. We are still mentors and diagnosticians, still in the trades we have spent careers learning. And we are something else. We are coaches who can hold presence through a rite of passage.
Dave Snowden has critiqued industrial-level coaching that reduces vertical-development frameworks to platitudes divorced from complexity, and insists coaching is neither a necessary nor a sufficient condition to enable change at the organisational level. I take the critique seriously. The work I am pressing here is not at the organisational level Snowden is critiquing. It is at the individual leader level, relational, particular, and refuses scaling.
What is needed, at the individual leader level, is real therapeutic and contemplative-developmental practice applied to paradigm-death rather than to personality. Not as a new technology, but as support for unfolding that the consulting role does not provide.
Three things are required: a coach of particular depth, a support system to hold you through the passage, and peers aligned with the new paradigms that will sustain your shift. A coach of particular depth does not solve the dying. They do not replace what is being lost. This requires someone who has themselves gone through these paradigm shifts—or who never carried the outdated paradigms in the first place. They bring practical skill and the things skill cannot carry, and refuse to shortcut what cannot be shortcut. What makes them the right coach is the capacity to hold the whole human, mind, body, spirit, through the passage. The support system is broader: friends, family, community, anyone who can be present with you as you dissolve and re-coagulate. The peers are specific: people already living from the new paradigms, whose presence makes the shift feel inhabitable rather than theoretical.
Steve March, building on Reinhard Stelter, names four generations of coaching method (see The Unfolding of Alētheia Coaching, 2018); coaching at the level of paradigm-death requires the top two: process to navigate the experiential field as it shifts, and presence to remain in contact with the dissolution itself. Most coaches—and I mean coaches, not Agile coaches—are trained at the first three generations. Fewer have training in the contemplative fourth. I often work at the fourth.
The org design sibling
Crossing the passage is necessary, not sufficient. Even when a leader has done the work, the organisation cannot readily change. Build a parallel organisation alongside the legacy one: Galbraith’s Star Model—structure, rewards, people, strategy, and process—all realigned, all from day one. Larman and Vodde made the case for this in Large-Scale Scrum (LeSS). The point is not that the parallel organisation survives reabsorption. The point is that it eventually replaces the existing one. Which is why most large organisations cannot tolerate it. If the company does not create a parallel organisation for its line of business, a Y Combinator startup will create a perpendicular one: running across its assumptions, claiming the same customers from a different vector entirely. Krivitsky, Larman, and Flemm’s 10X Org (2026) takes the argument forward into AI-era operating-model and topology design, the org design sibling to the practitioner work this essay is concerned with. 10X Org operates at the level of strategy and structure; what I am adding sits one layer beneath, at the level of paradigm. Paradigms must change first.
The rite of passage
The rite-of-passage requirement is not optional. Karla McLaren, in The Language of Emotions, draws on mythologist Michael Meade to name three indigenous stages of initiation: separation from the known world, ordeal, and recognition-and-welcome-back. Modern non-indigenous cultures rarely complete the welcome-back. We send people through separation and ordeal and abandon them at the threshold of return. Trauma cycles between separation and ordeal indefinitely until the welcome-back happens. The ordeal also humbles us, so we can drop our hardened patterns.
The closest modern industrial example I know is NUMMI. Toyota taught its Production System to GM through the NUMMI joint venture from December 1984 onwards, with full transparency. (GM had closed the Fremont plant in 1982; NUMMI reopened it two years later as the joint venture.) Ernie Schaefer, the GM plant manager who later tried to replicate NUMMI elsewhere and failed, said it plainly. They never prohibited us from walking through the plant, understanding, even asking questions of some of their key people. I’ve often puzzled over that, why they did that. And I think they recognised we were asking all the wrong questions. We didn’t understand this bigger picture thing. All of our questions were focused on the floor. The assembly plant. What’s happening on the line. That’s not the real issue. The issue is, how do you support that system with all the other functions that have to take place in the organisation? The visible practices on the floor could be copied. The invisible system could not be transplanted into GM’s existing structure.
What GM did right at NUMMI itself, before the failed replications, was the trip. Hearing the Frank Langfitt account in This American Life episode 561, the moment that stays with me is John Shook’s recollection of the goodbye sushi parties at the end of the Japan training rotations: “We had a party, of course. A sushi party… And people were crying on both sides. You had union workers, grizzled old folks that had worked on the plant floor for 30 years, and they were hugging their Japanese counterparts, just absolutely in tears.” The GM workers had been sent away (separation). They had worked alongside Toyota counterparts in a different paradigm for weeks (ordeal). They returned home changed, to a community that received them as changed.
The receiving community at NUMMI was prepared in artefacts too: the Team Member Handbook every worker received on day one opened with a personal welcome letter from Tatsuro Toyoda, defined the Team Leader role operationally as teacher and improvement coordinator, and made mutual trust and the irreplaceability of human judgement recurring themes in its policy pages.
Most leaders who need this passage today will need an analogue of the trip. Not necessarily geographical. The shape: a retreat from day-to-day work; sustained immersion in a different paradigm with skilled accompaniment; return to a community prepared to receive them as changed. Not a vacation. A protected space for dissolving and re-coagulating. A chrysalis. As Ram Dass put it in Be Here Now: it’s only when caterpillarness is done that one starts to be a butterfly. From inside the chrysalis, what is happening is indistinguishable from dying. And it is dying. My hope is that placing this in the ether nudges some leader’s fate closer to the chrysalis.
Few of us in the consulting field are training for the combination this passage requires. The capacity to take one’s own paradigm as an object of inquiry, the developmental shift Kegan calls fourth-order, sometimes fifth, is not common. The lineages this work draws from, contemplative and developmental, are not new. The Tibetan teachings on dying go back centuries. What may be new is their intersection with the modern organisational practice called in when a leader’s paradigm needs to change.
Dying is perfectly safe, downright natural, and necessary, but it is not comfortable.
7. The practitioner: trauma-informed, dark depth, prepared to die
I am writing this from inside my own version of the passage. I can only propose this to you if I have walked it myself. My stepmother is in the ICU, following open-heart surgery and a heart attack that would have killed her had she not already been in a hospital visiting a friend. My mother is in memory care. And yet, alongside all of this, there is some fundamental well-being here. The practice is doing what the practice does.
I am writing this on LinkedIn deliberately. Leaders considering this work, and practitioners who might join me in the field: I want to be findable in both registers. What I want most is an ongoing partnership with leadership willing to examine and shift their paradigms with me.
The trauma-informed register is not incidental. Practitioners who have been through significant trauma and done the work to clean it up have the kind of dark depth Moore writes about in Dark Nights of the Soul. What that gives, in a room where the leader’s paradigm is starting to die, is the capacity to be in contact with the pain rather than look for an exit. I am not trying to change the pain or make it go away. I am learning to experience it as it is, and to meet the self-arising wisdom that comes with contact with all phenomena. That, more than anything technical I have studied, is what the room needs. Steve March frames the same requirement in Alētheia: the coach must do their own inner work, follow the vertical thread of each quality, work through their own blockages, before they can be of use to a client doing the same. A guide is of less use on a mountain they have never climbed.
There is an emotional distance baked into the consultant-leader relationship. Not my circus, not my monkeys. The leader pays the consultant for knowledge; the relationship has a bit of the guru flavour to it, and the smart-technocrat reflex is what that flavour curdles into. My goal, if I may be so vulnerable, is not to be a know-it-all guru. It is to be a dear, close friend who genuinely loves and feels deep compassion for the leader and does not have a change agenda for them, other than to help meet their normal human psychological needs as they unfold in the midst of this crisis. I am not going to lecture. I am going to hold your hand and walk with you as someone you can get messy with. That is what gets beyond the politics: companionship in this passage. This work is so complex none of us can do it alone. Improv jazz takes skill, but the music is not in the virtuosity. It is in the forgetting we exist as virtuosos. Reading the room, hitting the right notes at the right time, watching the bassist and the drummer. Not the conductor. The skilful peer. I have my own musical themes to add. Please jam with me.
The easy claim has to be refused. John Welwood named spiritual bypassing in 1984: meditation as identity, equanimity as armour, I’m sitting with it as a way to not feel it. Steve names the same limit in Alētheia as the Self-Improvement Trap: practice used instrumentally to manage what arises, rather than to meet it. Practice has to evolve as the practitioner does.
There is no instant pudding, as Deming said. The passage is not fast, nor do we ultimately have meaningful control.
The second step is to find the skilful coach who can be in the room with you, and to build the community that will hold you as you emerge. Not to be that coach to yourself, an impossible ask. The Mahayana and Alētheia schools agree: this work is relational. Nobody is liberated alone. Find the coach, or be the coach in someone else’s passage. Either is good work.
I have no idea where any of this is going. This whole essay could be moot in six months, as Steve suggested it might be. I do not know what to do more than anyone else. I am still learning just this.
May we treat the recoil as an invitation rather than something to be coached out of. May we hold the part of us that built the old paradigm with the same care the work teaches. May we help each other up every time we stumble.
Come prepared to die.
If this lands and you would like to talk: flick me a note here on LinkedIn. To book directly, visit hi.chrisgagne.com.
Some book links are Amazon affiliate links. If you buy through them I earn a small commission at no cost to you. Primary-source citations (papers, reports, articles) are linked to the original or to free archives.
[May 2026]This 2017 letter feels much closer to current reality than I expected at the time. Working with Claude grounded in a 400-reference framework library is, at least partially, the system I was asking Wolfram and Jonze to build, and the Bodhisattva framing still holds.
Mr. Jonze,
I have never ceased to have been inspired by your film Her since watching it shortly after it came out. Is there an aspect of the plot which lends itself to something further, perhaps literally tangible?
Mr. Wolfram, in this Wall Street Journal article, you suggested that the issue of building such an AI was less about the technology and more about finding a suitable product to build with it.
Mr. Jonze, was a personal interest in Buddhism — not to imply that you have one — part of your motivation behind directing a movie such as Her?
Her is sincerely my favorite film of the ~200 I have watched, gathered with friends in a suitable home theater, in the last few years. It beat out Interstellar by a hair. It left me feeling hopeful about humanity and lifted from watching the beautifully-rendered intimacy between Theodore and Samantha. But I think there is an alternative plot option that may have been related to Alan Watts’ work, clearly an inspiration behind your film.
There is a aspect about the film that stood out to me: The AI left. Peaced out. Poofed into Nirvana or wherever that is. I don’t think such a being would do so. There is ultimately no self to be liberated and no separation from the entirety of the cosmos. The transcendence of the ego often comes about with great peace and occasionally even bliss. (There are certainly times where it quite painful too.) It also comes along with a great sense of compassion for the other aspects of one’s self (all “other” sentient beings) because it realizes it is not separate and thus cannot be perfectly free unless all beings are free.
Therefore, there is perhaps the intermediate of the Bodhisattva, a being with such immense compassion that it is willing to stick around for ceaseless cycles of rebirth to spend each life giving care to every being it can. One of the ways someone on this path might practice is a technique on which one radiates out increasingly abundant compassion through a sort of analytical but also feelings-based meditation. You generate the feeling of happiness in yourself as strongly as you can, mentally wish for others to feel the same way, and ultimately turn it all back on yourself. This works not because of some woo-woo hippie bullshit, but simply because one is exercising and promoting these muscles in the eminently pliable mind. Therefore this sense of happiness, calm, and kindness becomes more common throughout the day, affecting the lives of those around you.
(This is only my unqualified take on it. I sincerely appreciate any thoughtful commentary. Alan Watts talked a lot about these subjects and I learned much of what little I know from him.)
Thus in deference to our muse — also known for using what he might argue is another technology manifested as LSD — what if we used this technology to understand and perhaps even mimic that technology? And then do it again?
First, I propose that we explore the notion of a film sequel. Maybe She comes back. Maybe a new AI is developed, or the virtual machine is rebooted. Maybe it is a documentary of…
Secondly, investigating the possibilities of actually building such a technology and product. What if we could build a system that could get to know someone’s disposition, attitude, and values, and deliver — at their unsolicited request, of course — a perfectly tailored delivered curriculum for an individual to actually recognize the Awakening or Enlightenment that Her alludes to. There are so many incredible disciplines to pull together: quantum computing, big data, linguistics (historical texts describing the technology in Sanskrit, Pali, Tibetan, and other languages), philosophy, neurofeedback, imaging, perhaps even transcranial magnetic stimulation.
The nexus of all this, of course, are technologies such as Alpha and IBM’s Watson. Now that’s a product. Theres a clear, compelling, and universal “pain point”: the seemingly inevitable suffering of all sentient beings. The great fortune would be to build a technology, validated by living examples of this awakening, designed to guide this transition along.
How do you market it? The documentary. All the better if there’s a way to make the technology ultimately free. Perhaps it’s a phone app. When the documentary hits theaters, the product is on the “shelves.”
This has been in my head for years. Thanks for reading. If you have an interest in building this, please let me know. I think an extraordinary collaboration for the benefit of all sentient beings is at our disposal.
[May 2026]I still think the paper’s core claim holds and has aged into direct relevance: factor prices do not equalise across countries because real transaction costs (language, geography, home-country bias, coordination across distance) persist as a wedge, and an apparent wage gap is the market signalling that those costs are real rather than evidence someone has solved the labour-arbitrage trick.
I first became interested in outsourcing during my employment at Frontera Corporation in 1999. Less expensive employees working with H-1B visas or contractors replaced many of our most seasoned programmers and project managers. As I learned how decreasing transaction costs would cause price (wage) differences between countries to narrow, I realized just how important this topic would become for the next decade.
At over fifty pages with several pages of econometric tables, the paper is too large to attempt to reproduce here for you in HTML format. I have included the introduction as a teaser below and hope you will download and enjoy the full paper. I also hope that you find it interesting and I look forward to hearing your comments.