In The Timeless Way of Building, Christopher Alexander describes having strawberries for tea with a friend in Denmark. She slices them almost paper-thin. When he asks why she takes the extra time, she explains that this exposes more surface to taste. He writes: “I learned more about building in that one moment, than in ten years of building.”
For me, this is immediate, direct contact with what she’s doing. The way the fruit will be eaten guides the cut. It’s a small decision, but it helps me think about what we mean by taste when we’re making things with AI.
Alexander was trying to describe a quality that makes a place feel alive, with its parts belonging together and to the life lived there. He called it the “quality without a name” because words such as beautiful or comfortable each miss something. He also finds it in waves and a well-made fire. A fire can be too hot for comfort without losing that aliveness.
Think of a product that’s a joy to use, perhaps even a janky old motorcycle with worn paint and a few quirks. Something about the way it responds to you just… feels right. I think that can be a taste of the quality Alexander describes.
A product can look attractive and still be awkward to live with. Judging it involves paying attention to what happens when someone actually uses it, including things we might struggle to explain in a design brief.
The philosopher Eugene Gendlin describes how we can work with that difficulty. When searching for words, we can sense that a phrase doesn’t quite fit before we know what to say instead. He asks us to attend to a felt sense, the body’s sense of a whole situation, including more than we’ve already put into words. Trying another phrase can help us discover what we meant. In design, that might mean realising that the brief itself needs to change.
A language model can explain Alexander’s story and suggest cutting fruit that way. This doesn’t establish that it has tasted a strawberry, or that it can check its words against the bodily sense Gendlin describes.
A model could nevertheless learn useful distinctions from people who have those experiences. When we choose one design over another, we supply information even if we can’t fully explain the choice. I think a model could learn from those choices to propose things we value and help decide what to develop.
Perhaps some of what we call AI slop is work in which we feel this quality is missing. When we’re handed something that looks finished but leaves us struggling to make it useful, it can feel as though our time and circumstances weren’t worth much care.
Suppose a support team uses AI to draft courteous, accurate replies directing customers to another department. The departments close their tickets while the customer keeps having to explain the same unresolved problem to someone new.
If we’re reviewing only the replies, we might spend our time improving the tone. Following one customer’s attempt to get help could reveal that nobody has responsibility for resolving the whole problem. We might need to give one person responsibility for the case and authority to coordinate across departments.
The customer’s anger could help us notice this. Trying to make every exchange pleasant might mean soothing a justified complaint while leaving its cause untouched.
AI could help with that investigation as well as with writing replies. We could ask it to look across the case histories for repeated explanations or unresolved hand-offs, then check its findings with the people involved.
We could try a change on a recurring enquiry, then check whether people needed fewer repeat contacts and whether their problems were resolved. Staff and customers would need a way to tell us where the new arrangement still didn’t fit their needs.
Perhaps this fitness grows through contact: with the material, with the circumstances, and with the people who will live with what we make. We discover something we couldn’t fully specify beforehand, and allow it to change the work. Ideally, the people building the software encounter those situations directly: watch someone use it, hear what matters to them, and have the authority and supporting system to quickly respond to what they discover.
I’ve open-sourced grounded-forge, a working example of grounding an AI assistant in sources you trust. This post introduces it through one session, receipts included.
grounded-forge is my answer to the problem in the last post: an AI assistant defaults to the most-published take, not the best one, because the consultant layer outwrites the original thinkers and the models inherit the ratio. Ask a stock model about mission command and you get the LinkedIn consensus on mission command. The fix is to take source selection back.
The tool is open source. You pick the material you’d stake a real decision on. A model reads each source in full under a structured nine-pass protocol that checks every quotation verbatim against the text, then projects the source onto the tasks your assistant serves, once, at ingestion time. At runtime the assistant answers from that pre-audited material; the expensive, error-prone reshaping is already done.
What follows is the transcript of one session, working notes first, then the essay. I want to be plain about authorship: the transcript in the card below was generated by the model, and it matches the copy that ships in the repo; apart from trimming the lens list to the four the public repo ships, I have not touched it. The slate notes labelled “Chris” are mine, added afterwards to narrate what the system is doing at each step. My contribution otherwise is the question, the choice of sources, and the system underneath, which is built to ground every load-bearing claim in a source passage and to mark verified quotations as it goes.
The corpus behind the session is the demo library that ships with the repo (27 sources at the time), among them MCDP 1: Warfighting, the US Marine Corps’ capstone doctrine (public domain), and the OpenStax business textbooks, the standard management curriculum (openly licensed). My prompt was one line: “Please draw some interesting parallels between US Marines war-fighting doctrine and modern ‘textbook’ business administration.”
The [V] markers in the essay mean verbatim: checked against the ingested source text by a source-only audit at ingestion, before any question existed. (The auditor is itself a model working under protocol; no human read every line. What the marker buys you is a specific file and a specific passage to check when a claim looks wrong.) The trace at the bottom names which projections were read in full, and the closing section declines the one transfer the sources themselves don’t support. The transcript itself, without my notes, ships in the repo as a captured session. Where a [BT] appears, the idea is borrowed through the source from an author the corpus does not hold, and labelled as such.
Source-grounded synthesis · Protocol S (Synthesis) · 2026-07-23 · Claude Fable 5
USER · /answer-from-corpus
Please draw some interesting parallels between US Marines
war-fighting doctrine and modern “textbook” business administration.
ASSISTANT
I’ll answer this by drawing on the corpus. First, let me classify the
query shape and check for applicable lenses.
Shape: This is a Synthesis question
— it asks what the corpus says across two domains (Marine warfighting
doctrine, textbook business administration) and wants parallels drawn
between them. Breadth across sources is the point.
Chris: It has detected this is a synthesis query. That is one of three shapes it routes: a named lookup reads one source and stops; a diagnostic maps your situation into one task domain; synthesis wants breadth across the corpus. The shape decides which indexes get read and in what order.
Lens-applicability check: Lens-neutral. The query
asks for an intellectual synthesis between two doctrines — not a
role-bound artefact, not a reader-typed deliverable, and not an
application of the named-person lens. No lens applies.
Chris: It has decided not to apply a lens, as none fits the question. A lens reweights what a distillation surfaces for a particular role (a CTO and a builder need different salience from the same source), and the check runs before decomposition, while it can still shape the sub-claims. It declines the four on offer rather than forcing one.
Sub-claims (the parallels to test):
Nature of the decision environment — friction and uncertainty vs
bounded rationality and scarcity
The decision cycle — OODA vs the textbook multi-step process
Who decides — mission command vs
centralisation/decentralisation
Priority-setting — Schwerpunkt vs the textbook prioritisation
apparatus
Error asymmetry — inaction vs commission
Doctrine as thinking vs framework as tool
Limits of transfer — the adversarial Zweikampf frame
Tool calls – Read slug-table.json
(full) — 27 sources; 00q = mcdp1-warfighting, plus the
OpenStax family. – Read concept-index.json (full; exceeded
the single-read cap — two paginated reads, 4,210 lines). – Read
distillations/decision-making/task-index.json (full; two
paginated reads, 1,809 lines).
Chris: This is the routing pass, and the reads are whole-file by rule. The concept index is bigger than a single read allows, so it paginates to the end rather than grepping for a likely key. Grep answers the question you thought to ask; a full read is what lets the index volunteer sources you didn’t think to ask for, which is exactly what happens next.
Routing findings. The concept index carries a dense
MCDP-1 concept family (ooda-loop, mission-tactics-auftragstaktik,
intent-two-levels-up, schwerpunkt-main-effort,
leniency-overbold-severity-inaction, zweikampf-violent-struggle) and
routes the business side beyond the management shelf:
theory-of-constraints and
relevant-costs/sunk-costs live in 00b
(Principles of Accounting Vol 2), and sunk-cost,
opportunity-cost, and marginal-analysis concepts live in
00e (Principles of Economics 3e). One gap to
report honestly: the task index carries no rows for
00q at all — the situation router does not yet
cover the corpus’s newest source. Routing for the Marine side therefore
rests on the concept index and slug-table, which is sufficient here.
Chris: The payoff of the full reads is in the findings above: accounting and economics surfaced for a Marines question, which no skim of filenames would suggest. And when the router itself had a hole (no task-index rows for the Warfighting source), the session reported the gap instead of papering over it. That is the no-silent-degradation rule doing its job; I fixed the index the same day.
Candidate set and dedup.00q
(Warfighting), 00j (Principles of Management),
00h (Organizational Behavior), 00g
(Introduction to Business), 00b (Accounting Vol 2),
00e (Economics 3e). 00c (Business Ethics)
matched on stakeholder-theory for sub-claim 7, but 00h
(Follett) and 00j (Freeman) already carry that sub-claim —
skipped under the diversity-aware dedup rule.
Every sub-claim now has the Warfighting side and at least two
textbook-side sources, with convergence and divergence both represented.
The accounting and economics rows extend the comparison well beyond the
management shelf. Let me write the essay.
Chris: Six distillations read in full, not as retrieved chunks: the four any skim would find, plus the accounting and economics volumes the concept index surfaced. Each is a source already projected onto decision-making at ingestion, quotes checked and marked in place. The costly work ran once, before any question existed. It also declined a seventh read: Business Ethics matched, but the stopping rule found the sub-claim already carried.
Two Doctrines of
Deciding Under Uncertainty
Read MCDP 1: Warfighting against the OpenStax business shelf
— not just the management row, but the accounting and economics rows
beside it — and the same question keeps surfacing from both sides:
how do you decide well when you can’t see the whole board? The
parallels are real and useful. So are the divergences. And one claim the
narrower comparison invites turns out to be wrong once the whole shelf
is on the table.
Where they agree:
the constraint is permanent
MCDP 1 names the environment before naming any method: war is
friction, uncertainty, fluidity, disorder — constitutive features, not
exceptions. The doctrine is blunt that “no degree of technological
development or scientific calculation will diminish the human dimension
in war” [V]. Decision frameworks that assume the world is reducible will
fail, because the world isn’t.
The management texts arrive at the same place through cognitive
science. Bounded rationality anchors both OpenStax management volumes:
“for complex issues we cannot be completely rational because we cannot
fully grasp all the possible alternatives, nor can we understand all the
implications of every possible alternative” [V]. And the economics text
arrives through a third door, defining its whole subject as “the study
of how humans make decisions in the face of scarcity” [V] — adding that
“the ultimate scarce resource is time — everyone, rich or poor, has just
24 expendable hours in the day” [V]. Three disciplines, one founding
condition: the limit on the decider is permanent, so you decide inside
it. Introduction to Business concedes the same about control:
“managers are primarily adapters to, rather than agents of, change” [V].
Neither tradition believes the decider owns the field.
The decision
cycle: same skeleton, different clock
MCDP 1 credits Boyd for Observe–Orient–Decide–Act and stakes its
claim on tempo: the side that completes the cycle faster, or
that the adversary cannot orient to, wins. The textbooks run the same
loop at lower velocity — recognise, generate alternatives, analyse,
select, implement, evaluate in the management volumes; a five-step
variant closing with feedback in Introduction to Business,
which warns that managers can get dangerously stuck at any stage.
Analysis paralysis is a slow OODA loop by another name.
The first-order emphasis gap is real: the management version
optimises the quality of each step, the Marine version the
speed of the whole loop relative to the opponent’s. But the
economics row complicates the neat contrast, because it prices delay
explicitly: “the opportunity costs of waiting time can be just as
important as costs that involve direct spending” [V]. The textbook shelf
does carry a clock; it lives in the economics text rather than the
management one, and it prices your waiting against your own alternatives
rather than against an adversary’s cycle.
Who decides:
mission command in three vocabularies
MCDP 1’s answer to who should decide what is mission
tactics: the senior gives the what and the why — the
commander’s intent — never the how, and intent must be understood
two levels up. Decentralisation works only when subordinates
hold the intent; without it you get incoherence, not empowerment.
Introduction to Business debates the same trade-off as
centralisation vs decentralisation — warning both that
centralisation can prevent quick local decisions in dynamic environments
and that decentralisation without skills or training can produce costly
mistakes — and renders it structurally as organic versus mechanistic
design. The accounting volume adds a third vocabulary the narrower read
missed: responsibility centres, which align decision
authority with information access and accountability. That is the
information-logic of mission command in accounting dress: push the
decision to where the information lives, and hold the decider
accountable for what they control. What the textbook shelf still lacks
is MCDP 1’s sharpest instrument — intent two levels up as the
specific content the empowered subordinate must hold. The
textbooks say decentralisation needs skills; the Marines say precisely
which skill: the boss’s boss’s purpose.
Priorities:
the textbooks have a Schwerpunkt after all
Here the whole shelf corrects the essay a narrower read produces.
Compare MCDP 1’s Schwerpunkt — name one main effort;
everything else supports it; supporting yields when they conflict — with
the management shelf alone, and the textbooks look like
better-prioritisation people: Drucker’s eight goal areas, SWOT, weighted
analysis. Prioritising better preserves the multi-priority frame; naming
a main effort breaks it.
But managerial accounting carries the textbook tradition’s own
one-filter discipline: constrained-resource allocation.
When a resource binds, rank every product by contribution margin per
unit of the constraining resource — not by unit margin. The
highest-margin product is often the wrong priority; what matters is
yield against the single binding constraint. The volume carries this
from Goldratt’s Theory of Constraints [BT], and it does to product-mix
decisions exactly what Schwerpunkt does to operations: one thing
governs, everything else is ranked against it.
The residual difference is sharper for being narrower. The Marine
main effort is chosen — a bid, an act of will that can be
shifted mid-fight as gaps appear. The accounting constraint is
discovered — an empirical fact about capacity you find and
exploit. One tradition names its focus; the other finds it. That is the
true divergence, and it is smaller than the one the management shelf
alone suggests.
The asymmetry
on inaction: the mirror has a crack
MCDP 1 rejects zero-defects outright: severity on errors of
inaction, leniency on overbold errors. The management texts brake
in the opposite direction, and for cause: their catalogued failures are
failures of commission — escalation of commitment, “the tendency of
decision makers to remain committed to poor decision, even when doing so
leads to increasingly negative outcomes” [V], groupthink [V], and the
remedies are devil’s advocates and precommitted monitoring.
Yet the economics row again refuses the clean mirror, because it
carries anti-hesitation machinery of its own. It names why waiting feels
safe: loss aversion — a loss pains roughly 2.25 times more than an
equivalent gain helps [BT] — and it orders the release: “the lesson of
sunk costs is to forget about the money and time that is irretrievably
gone and instead to focus on the marginal costs and benefits of current
and future options” [V]. The accounting volume drums the same refrain
into every short-term decision category. So the honest contrast is
institutional, not intellectual: the textbook shelf hands the
manager tools against both failure directions and leaves the
calibration to judgment; the Marines legislate the asymmetry as
policy, because their characteristic death is hesitation. The tools
exist on both sides. Only one side turned them into doctrine.
Doctrine as thinking:
three-way agreement
MCDP 1’s most general claim — doctrine is a way of thinking, not a
checklist — turns out to be the one the whole shelf endorses. The
management texts warn against running frameworks mechanically and insist
a good process can yield a bad outcome under uncertainty. The accounting
volume makes the same move in its own material: costs are not
intrinsically fixed or variable — the classification is
decision-relative, and the practitioner’s question is “fixed for
what decision?” A label that must be re-derived per decision is
judgment wearing a taxonomy. The framework is scaffolding; judgment is
the thing. On this, no daylight.
The one parallel that doesn’t
hold
MCDP 1 frames war as Zweikampf — a violent struggle between
two hostile, irreconcilable wills — and its edge follows from that
frame: tempo, surfaces-and-gaps, getting inside the opponent’s cycle.
The Warfighting distillation itself flags this as the piece that usually
does not transfer. The textbook shelf encodes the opposite cosmology
three times over. The management row offers Follett’s integration — both
parties reaching the creative agreement that satisfies both [BT] — as
the strongest conflict outcome, and flatly rejects harming stakeholders
for shareholder return: “increasing the wealth of shareholders is not an
acceptable reason for causing harm to others” [V]. And the economics row
is built on the founding scene of mutually beneficial exchange:
comparative advantage says the counterparty is a trading partner whose
different opportunity costs are the very reason to deal, not a will to
break.
So the synthesis holds, sharpened: the machinery transfers — OODA,
mission command, constraint-governed priority, the case against
hesitation, doctrine-as-judgment — and much of it has closer textbook
cousins than a narrow read reveals. The cosmology does not transfer.
Reach for the Marine playbook when there is genuinely an adversary and a
clock; reach for the shelf when the room holds stakeholders and trading
partners, which for most of management is most of the time. The mistake
that survives every widening of the corpus is the same one: using the
Zweikampf frame on people who were never your enemy.
Chris: The trace is the receipt: protocol, corpus, the lens ruling, which sources fed each sub-claim, what was read in full, and even the read it declined. It is what lets you audit the answer instead of taking its fluency on trust. A trace that names what was read shows what wasn’t.
That’s the whole exchange, trace and all. The build system, the 9-pass ingestion protocol, and the evaluation showing where the architecture wins, and where a naive corpus read beats it, are open at github.com/chrisgagne/grounded-forge.
If your work leans on sources you actually trust (doctrine, standards, your own case notes), the pattern transfers: each source is read once under audit and projected onto the tasks you repeat, and answer time becomes a lookup. I run the same machinery for After-Action Review and retrospective facilitation. My hunch is that the audit trail matters more than the essay. Fluent prose from a model is cheap now; what I wanted was prose I can check.
I first published a version of this piece on LinkedIn in 2021, when the argument ran on Weinberg’s table and Kniberg’s prioritisation illustrations alone. They carry the “don’t do everything at once” point well enough. What I have built since is a working cost-of-delay tool, grounded in Don Reinertsen’s Principles of Product Development Flow—the source of both the cost-of-delay economics and the WSJF / CD3 prioritisation heuristic—and the practical cost-of-delay work of Joshua Arnold, who did as much as anyone to turn it into something teams actually use. The “cost of delay” is no longer a rhetorical lever for me: the tool ranks every story by it, and where the business states what the work is worth, it can put an estimated dollar figure on the delay. Also, the Cost of Delay work has shown me that most of the value comes from focus, not necessarily getting the prioritisation perfect.
Gerald Weinberg’s book “Quality Software Management: Systems Thinking” is more than 30 years old. While it’s not one of the most highly-read and recommended classics in the adaptive canon, it’s the source of a frequently quoted table of data:
Visualised as a graph, the waste caused by context switching really stands out:
Weinberg’s figures, Table 2-1. Each coloured block is one project’s share; grey is time lost to switching.
Worse, the waste caused by project switching isn’t due only to the losses related to cognitive overhead. Failure to prioritise often leads to less revenue and even building the wrong product.
(Credit where credit is due: I first saw a version of the following illustrations presented by Jeff Sutherland, who adapted it from Henrik Kniberg. I’ve created new illustrations and expanded on them a bit. Kniberg’s own video on this exact topic is very much worth a watch.)
Let’s suppose your company wishes to ship three products: A, B, and C. To ship a product, a development team must complete tasks 1, 2, and 3 corresponding to that product:
Most companies are not very good at prioritising work. As a consequence, the prevailing belief is: “Everything is important. Get started on everything immediately!” The traditional delivery timeline might look like this:
A blue tag marks the day a product ships.
Depending on your software and how you break down your work, your roadmap probably looks like:
A blue tag marks the day a product ships.
Either way, notice that we are interleaving each task and that products A, B, and C are ready at roughly the same time, several months after we started.
The adaptive approach is very different. Instead, we proactively prioritise the work and focus on limiting our work in process:
A blue tag marks the day a product ships.
There are at least three significant advantages to this adaptive approach: lower cost, more revenue, and better product/market fit.
Lower cost of development
We lose 20% or more of our productivity in the traditional approach due to context switching waste. In this example, the company could go about twice as quickly if they switched to developing one product at a time instead of three.
Grey: time lost to switching between products.
More revenue, earlier
The switching waste is the smaller loss. Because we have nearly finished products B and C before finishing A, we had to do nearly three times the work (not including the context-switching waste) before we could ship Product A (in late May). Had we prioritised and focused, we would have been able to ship Product A in early February. Had we done so, we may have been able to collect revenue and feedback from our customers starting nearly four months earlier. In fact, the savings from not context-switching between projects may mean that we could ship products A, B, and C before we would have been able to ship just A in the traditional model.
Hatching: a product waiting between its tasks. Blue: a shipped product earning.
Better product/market fit
Products are rarely independent, and building the wrong product can be costly. In this example, we believed before starting the work that the customer needed Product C. We took the adaptive approach and built A (more valuable and/or cheaper) first. When we delivered A to our customers in early February, we learned that they liked it a lot and didn’t need C after all. Instead, they wanted us to work on D. We took February to finish most of B and get the initial prototypes for D built.
Dashed: planned, never built. Blue: a shipped product earning.
We can see that prioritisation and focus serve everyone: developers, stakeholders, and customers.
Good prioritisation is not a zero-sum game
Further, good prioritisation is usually not a zero-sum game. Many stakeholders will argue aggressively for their product to be worked on right away (and it’s no surprise: they may even have a bonus tied to it being completed by a certain date). However, the waste caused by project switching and the cost of delay is so significant that failure to prioritise may make all stakeholders worse off. In this case, the stakeholders behind projects A, B, and C agreed to prioritise based on the cost of delay and all were better off than if they had insisted that everyone’s work be tackled at once.
The same three epics, priced
Five years on, this argument became software. The infographic below, from my Delivery Intelligence work, prices the same scenario with cost of delay: keeping all three epics in flight collects $42.50 of value by the end of July; finishing one at a time collects $138.75 in the worst order and $176.25 in the best. In this toy scenario that is roughly 3–4× the value, and the ordering mattered far less than the focus.
Each epic needs three large tasks. Once shipped, A earns $10 a month, B $5 and C $20. Interleaved, each task runs about 60% longer: at three projects in flight, roughly 40% of the time is lost to switching. Relatively speaking, the order made little difference; the value came from the focus. Interleaving against one-piece flow, after Henrik Kniberg.
Feel it in five minutes: the character factory game
You don’t have to take my word for any of this, or Weinberg’s. Grab a pen, a sheet of paper with three columns, and a one-minute timer. Three customers each want a complete set of 25 characters—letters, numbers, vowels—and nothing partial counts. Round 1: keep all three orders moving by writing one character for each customer in turn. Round 2: finish one customer’s set before starting the next. One minute each. (The game is Henrik Kniberg’s multitasking name game, with a modification I learned from Monica Yap and a further modification of my own.)
The card assumes the same writing speed in both rounds. Focus still wins—two complete sets against none—because a complete set is what a customer can actually use. Run it live and a second effect appears: your total character count drops in round 1, because every switch costs a beat. That drop is Weinberg’s table on one sheet of paper, and the delivered sets are the revenue argument in miniature. Try it with your team this week.
Some book links here are Amazon affiliate links; if you buy through them I may earn a small commission, at no cost to you.
Chris Gagné built grounded-forge, an open tool that grounds an AI assistant in sources you trust.
Your AI assistant doesn’t hand you the best thinking available. It defaults to the most-published.
The mechanism is a corollary to Larman’s Laws of Organizational Behavior. Original thinkers write slowly and rarely. Displaced managers become coaches, coaches become consultants, consultants become prolific content producers, and that derivative layer outwrites the primary sources by something like an order of magnitude. Models train on the written record, so they inherit the ratio. Ask a general-purpose model how to run an incident review or structure a hard decision, and the default answer is the consultant-frequency consensus, not the sources you’d stake a real decision on. Web search doesn’t fix it by itself. It samples the same distribution: more language, not more truth.
I’ve open-sourced grounded-forge, my working answer to that problem. It’s also a hands-on example of the principle I build everything around: put the LLM where it earns its place, and stay deterministic where consistency, auditability, and cost win.
What it does
You pick the sources you trust. An LLM reads each one in full, not chunked, under a structured nine-pass protocol that produces verbatim-cited references and pre-projects each source onto every task your assistant serves: decision-making, stakeholder engagement, after-action review. A source-only audit checks every claim against the text before anything ships.
At runtime the model doesn’t re-project raw chunks onto your task. It picks a pre-projected, pre-audited artefact and answers from that bounded, cited material. The design shrinks the room to hallucinate, and the source-only audit that gates every claim before it ships catches much of what would otherwise slip through. Not zero. That it beats standard retrieval is argued from the design, not measured. You pay for the synthesis once, under audit, and every query after that is selection rather than re-derivation.
That’s the whole bet. The expensive, error-prone work is reading the source and projecting it onto your task. Standard retrieval does that again on every query, from raw chunks, with the model reshaping fragments under time pressure. grounded-forge moves it to ingestion time, does it once under discipline, and reads the result at runtime.
Where it wins, and where it doesn’t
The evaluation results are in the repo, including where the architecture loses.
On canonical public material with clean filenames, a naive corpus read wins and the matrix is pure overhead. In the evaluation, tidy author-and-topic filenames gave the naive read enough routing signal to win outright. The likely reason is that the name keys a strong training prior, so the model routes in one read and spends the rest on content, while the projection machinery re-derives what the prior already holds. I ship that as a bounded claim, not a blanket one.
The matrix earns its keep where your organisation’s real knowledge lives: engagement documents, incident histories, your own frameworks, the material no training run has seen and no model can route without help. There the training prior has nothing to invoke, and the curation does the work. A second result pointed the same way. The lens layer, which reweights a distillation through a role or a methodology stance, beat its lensless variant by 0.6 rubric points on the same non-public prompt, in the direction the lens was designed to move. Directional, on a small sample, and reported as such.
One honest caveat on the audit numbers. A fresh-context audit by a different model family found 97.45% of 9,364 deep-reference claims free of hard errors before repair; the repairs also covered minor drift, such as dropped qualifiers. It’s a discipline check, not a truth oracle. It makes the assistant’s fidelity to the chosen source inspectable; it doesn’t establish that the source is right about the world. That judgement stays yours.
For builders
The architecture is a reference × task matrix. One axis is your sources. The other is the tasks your assistant actually does. Each cell is a distillation: one source projected onto one task, carrying verbatim quotes and evidence markers in-band, so the load-bearing claims arrive already cited. Retrieval becomes a lookup into that grid rather than a reshape of chunks.
The demo corpus ships 28 sources across five task axes, 119 distillations in all. It isn’t 28 × 5: some source-task pairs are deliberately skipped where the projection would add nothing the matrix doesn’t already carry elsewhere. Those 119 cells compile into five distributable assistants from one build command. Two of them, decision and stakeholder, are built from the identical source library and the identical runtime indexes; they differ only in their task distillation directory and the CLAUDE.md that briefs the assistant for that task. diff -rq between the two apps returns exactly those three lines. That’s the matrix made concrete: one corpus, many task readings, each shipped as its own assistant.
Fork it, strip the demo content in one command, and ingest your own corpus. The reference tier is the audit-of-record and stays at corpus level; the compiled app ships only the distillations and the routing indexes it needs. It runs in Claude Code today. On licensing, the substrate (build system, scripts, skills) is MIT, the long-form prose I wrote is CC BY 4.0, and each source-derived reference inherits its own source’s licence, stamped in the file’s frontmatter. The corpus is plain markdown throughout, so an operator who walks away takes their files and goes. Nothing in the design locks the corpus to me or to any one vendor.
Ingestion runs with a 9-pass protocol: here’s the tl;dr:
Pass I checks the deep reference against the converted source text in a fresh context.
The organisational work
Many AI programmes start at the coder-tooling layer and stop there. Roll out a chatbot, wire up autocomplete, call it done. The harder and more valuable work sits above that layer. It’s deciding what your organisation treats as authoritative, and building the machinery that holds the line when a model would otherwise reach for the consultant-frequency mean.
Pick your sources. Project them onto your work. Make the assistant read through your windows, not the training distribution’s.
You shipped more this quarter than last. Your customers are no better off. Which number is the work actually being judged on?
By now you know the move: find the axis where your claimed pole and your lived pole pull apart, then give yourself to that analysis. If output-versus-outcome is already the tension eating your quarter, this is your chapter. If it isn’t, read it lightly and save your energy for the axis that is.
Subordinating non-constraints was chapter 8’s move. It pays off only when the constraint serves the right number. Output and outcome are different numbers, and the gap between them is where feature factories live.
🎧 Prefer to listen? This chapter is narrated in my own voice with ElevenLabs on Spotify (~18 minutes).
The two poles
Pole A: output. The unit of management is velocity, units shipped, features released, work completed.
Pole B: outcome. The unit of management is customer progress, market position, mission impact.
This axis is well known among engineering leaders. The vocabulary—feature factory, output vs outcome, customer progress—has been in circulation for well over a decade.
Where Pole A is right
Pole A is right when output and outcome have been shown to track each other tightly in the relevant domain. In a sales organisation with a mature product, deals booked may convert predictably to revenue over months. In a fulfilment operation, units shipped may equal customer demand met, provided return rates remain low and margins stable, and in a regulatory pipeline, filings submitted may track approvals secured when the regulator’s response time is well-characterised over multiple cycles.
The metric I was judged on was velocity. Wherever it became the watched number, estimates grew to meet it; I once saw a team carrying a velocity in the hundreds, the number climbing while nothing shipped any faster. Metrics like these earned their place first: deals booked, units shipped, and features released were proxies that tracked the customer outcome closely enough that managing the proxy amounted to managing the outcome.
None of that discipline was wasted. Counting output is part of what makes anything ship at all. A team that can’t count what it has produced can’t tell learning from activity.
Pole B doesn’t abandon counting. It asks you to notice when output stops tracking the outcome. Output was a clean proxy until the conditions that made it clean stopped holding; after that, counting output means counting the wrong thing.
Where Pole B is right
Pole B is right whenever the firm produces more output and gets worse outcomes. This is the empirical pattern in feature-factory product organisations. The velocity’s high and the customers still leave. At one engagement, the teams had been quick before the change, but they were building the wrong product quickly. Afterwards they were slower, and the quality was better at the end. Pole B costs the leader something real: standing in front of a falling output number on purpose.
Clayton Christensen’s Jobs to Be Done theory, laid out in Competing Against Luck, gives the central reframe. The customer isn’t buying your feature. The customer is hiring your product to make progress on a job.
The classic example is the milkshake a fast-food chain sold in surprising numbers before 9 a.m. to commuters buying nothing else. Commuters hired the milkshake to keep a long, boring drive interesting, one-handed, and to hold off the mid-morning hunger. The milkshake was hired because it outlasted every competitor—a thick shake takes 20 minutes through a thin straw—and the cup fit the cup-holder.
Improving the milkshake meant making it last longer, not making it taste better. The Pole A move (survey the customers, ask what flavour they want) produced flavour iterations that didn’t move sales. The Pole B move (understand the job) produced an actionable insight.
The chain sold the milkshake in surprising numbers before 9 a.m., to commuters buying nothing else. After Clayton Christensen, Competing Against Luck.
Marquet: red work and blue work
David Marquet’s Leadership Is Language names an operational discipline that fits the output-outcome distinction. Red work is execution under pressure: doing what has been decided. Blue work is judgement and pause: deciding what to do.
The Industrial-Age inheritance divided planning from doing. A few people decided; everyone else executed. Marquet recasts those activities as blue work and red work, then argues that work now requires everyone to do both.
Proving and performing give way to improving and learning.
The Pole A leader who manages to features-shipped runs in pure red mode. The team decomposes the goal into tasks, executes them, and reports completion. Blue work has no place to happen because the language of the work has no opening for judgement.
The Pole B leader inserts blue work into the cadence. The team pauses, checks whether the output delivered the outcome it was meant to serve, and changes the next round of red work based on what it learned.
We committed to ship X by Q3 closes the loop on red work. We hypothesised that shipping X would produce outcome Y by Q3; what did we learn? opens the loop into blue work.
Marquet’s own protocol on the USS Santa Fe, replacing “request permission to…” with “I intend to…”, moves the judgement to the person closest to the work; chapter 17 runs it in full. What matters here is what the leader states, what military doctrine calls commander’s intent: a plan that specifies the steps dies the moment events outrun it, and the statement that survives is the goal and the end-state, compact enough that people down the line can improvise toward it.
Innovation accounting
Eric Ries hands the output-counting leader an accounting discipline for Pole B. Vanity metrics rise regardless of whether anything is working; actionable metrics change what the leader does next.
Here’s the test I use to tell them apart. Pick the metric. Predict what it will do over the next two weeks or months. Look at it. Did the prediction match? If it didn’t, what would you do differently?
The chapter’s own test; vanity and actionable metrics after Eric Ries, The Lean Startup.
If the answer to that last question is nothing, the metric is vanity: you’re not running on data, you’re running on its appearance.
Innovation accounting, Ries’s term for tracking learning per unit of investment, converts Pole B into language the Pole A board can use without losing accountability. It is the chapter-7 board script in a different currency: commit to learning a named thing, by a date, at a cost, then report what was learned and what to test next.
Cagan: valuable, usable, feasible, viable
Marty Cagan’s outcome test is the simplest I have seen for product work. A feature is worth shipping if it is valuable, usable, feasible, and viable.
Valuable means it produces an outcome the customer cares about. Usable means the customer can access that value. Feasible means the team can build and operate it within the firm’s constraints. Viable means the business itself can sustain it.
The failure modes are easy to recognise. A Pole A team may optimise for feasible, building what it can build and shipping features that are technically clean but irrelevant to customers. The most user-research-heavy teams sometimes over-rotate on valuable and ship features the team can’t operate at scale.
A feature is worth shipping only if it lands all four. Run the four-part test on last quarter’s releases and count how many clear all four.
Edmondson’s intelligent failure
Fail fast becomes operational only when a failure meets Edmondson’s four criteria for an intelligent failure, which chapter 10 walks in full. Until then it is a slogan a Pole A leader can adopt while keeping the old metrics: failed experiments per quarter is an output number too, and it passes the vanity test no better than features shipped.
A representative case
Picture a SaaS product team that has been measuring features shipped per quarter for two years. The number goes up steadily. The customer churn rate also goes up steadily.
The CEO can’t reconcile the two numbers. Engineering leadership points at the customer success function; the customer success function points at product. The board is being told we are shipping faster than ever in the same quarter that renewals have started falling.
The product had grown surface area faster than its customers could absorb it. A representative composite from the chapter, drawn to schematic scale: two years of a SaaS product team counting features shipped per quarter. Not data.
The Pole B move arrives through a customer advisory board. Several customers on the board, on separate calls, name the same pattern: the product has grown surface area faster than they can absorb it. The features they asked for two years ago are buried under features they didn’t ask for, and the team they originally trained on the platform has churned. The outcome is being eaten by their accumulation.
The retrospective surfaces the structural problem. Every feature shipped had been valuable, usable, feasible, and viable at the time it was scoped. Nobody had re-evaluated it against the integrated product experience.
The fix wasn’t to abandon shipping. The team introduced a quarterly outcome review alongside the existing roadmap review and asked, for each major feature shipped over the past four quarters, whether customers were using it, whether that use produced the outcome the feature was supposed to deliver, and whether the maintenance cost still tracked the value.
Half the features failed the test. A third were deprecated. The next quarter’s roadmap was leaner, and the cohort that had been churning steadied.
The engineering leader running the original output count wasn’t a bad leader. The previous CEO had measured them on output. The investor decks had measured them on output. The internal performance reviews had measured them on output. The system around the leader had reinforced Pole A for years.
So name what the discipline got the team—predictable cadence, investor confidence, a high-trust engineering culture—and then name what the new conditions required. The leader didn’t need a different character. The leader needed a different measurement system, and time to adjust to it.
The diagnostic move
Three questions for last Tuesday’s launch celebration.
Which pole was I claiming? In the language I used about the launch, did I name what the customer would now be able to do, or did I name what the team had shipped?
Which pole did the team’s behaviour show? Look at where the team’s attention went in the week after the launch. Did it go to customer use or to the next item on the roadmap?
Which pole does the work actually require? If output and outcome track tightly in your domain, Pole A is legitimate. If you are shipping more and the customer outcomes are flat or worse, Pole A has stopped being a proxy and started being a substitute.
The board deck often locks in this distinction. Whether the CEO or CTO owns the page, the metric used to describe engineering shapes incentives below. If you run engineering, the metric your CEO uses to describe your function to the board governs the output-versus-outcome question.
Put the metric that currently judges your function beside the metric that tracks customer progress. In most software-dependent organisations, these are different numbers. A CEO who describes engineering to the board in features-shipped vocabulary creates Pole A incentives three layers down, usually without intending to.
Ask for one sentence in the next board update that names what customers can now do, not what the team shipped. That is the Pole B move at board level. The team will read the update.
Give your CEO this sentence: “Customers are now doing something they couldn’t last quarter; the features we shipped are how we got there.”
The exercise
Rewrite a recent feature as a job to be done. Christensen’s test is specific: a job is the progress a person is trying to make in a particular circumstance; state it in verbs and nouns, and check that candidates for the job come from different product categories.
The job-story template—“when [situation], I want to [motivation], so I can [expected outcome]”—is a usable scaffold. It comes from the job-story tradition, not from Christensen. The format forces the team to name what the customer is hiring the feature for.
Notice how often the feature solves the firm’s job rather than the customer’s. When my quarterly performance review is approaching, I want to ship a visible feature, so I can demonstrate productivity to my manager is a job. It belongs to the employee, not the customer.
Run a companion vanity-vs-actionable metric audit. List your top five metrics. Mark each as vanity (it goes up regardless of whether anything is working) or actionable (it changes behaviour when you look at it). Defend each judgement to the team.
The usual engineering ones rarely survive it. Velocity moves whenever estimates move, and stories shipped is the same count in different units. Utilisation is worse than either: it rises as queues lengthen, so a high number is as likely to be a warning as a win.
If leader-follower dynamics turn out to be your live axis, run Marquet’s “I intend to…” protocol for a week. Chapter 17 lays out the full version. Even one week of replacing “request permission to…” and “recommend that we…” with “I intend to…” shifts who carries the cognitive load of the decision; run it and watch where decisions land.
Also touched: Eric Ries, The Lean Startup, on innovation accounting and vanity metrics; Douglas Hubbard, How to Measure Anything, on the test that a measurement only earns its keep if it changes a decision; Marty Cagan, Inspired, on valuable / usable / feasible / viable; Jeff Patton, User Story Mapping, and the talk Output vs Outcome & Impact (11 min, watch); Amy Edmondson, Right Kind of Wrong, on the four criteria for intelligent failure (developed in chapter 10).
Go deeper: Chip and Dan Heath, Made to Stick, on commander’s intent in organisational-messaging terms; USMC, MCDP-1 Warfighting, Ch 4, on mission tactics and commander’s intent, with the German antecedent Auftragstaktik from nineteenth-century Prussian military reform; Frederick Taylor, The Principles of Scientific Management (1911), on the planning/doing division Marquet recasts as blue and red work; Alan Klement, Replacing the User Story with the Job Story (2013), on the job-story form, which originated at Intercom.
Manage to outcomes and some of them will miss. Watch what you defend when yours does. The call you made was a testable bet; the thing you protect in the review is usually your identity, and the second job eats the first. Chapter 10 sits with that review, the one after a decision fails, and asks what the miss costs you next time: whether you spend it shielding the person who made the call or finding out what the call was worth.
I work with engineering leaders on exactly this kind of paradigm work, the deeper the better. If it’s live for you, I’m happy to talk: schedule a 30-minute virtual coffee at hi.chrisgagne.com.
Some book links here are Amazon affiliate links; if you buy through them I may earn a small commission, at no cost to you.
[May 2026]I now reach for Schein on culture, Westrum on generative-vs-pathological organisations, Edmondson on psychological safety, Senge on learning-organisation discipline, and Galbraith’s Star Model alongside Org Topologies for structure as the more rigorous current vocabulary for what this post calls “structure and culture.” The doing-vs-being distinction still holds; the toolkit has matured. I’ve also since added a fifth box at the cheap end of the diagram, to the left of tools: “terms,” changing the words you use. It’s the easiest change and the emptiest one, and I can’t claim much originality for it. Craig Larman named this years ago in the second of his Laws of Organisational Behaviour: any change initiative gets reduced to redefining or overloading the new terminology to mean basically the same as the status quo. New words, same operating model.
[August 2026]The NUMMI listening questions I set out here, about what changed in the plant and what GM could not replicate elsewhere, are the ones I work through at length in Chapter 3 of Come Prepared to Die. Same plant, read the other way round: the chapter takes NUMMI as the inverse case, where GM copied the production system and left the structure alone. AI now hands you the tools-and-process win almost for free, which leaves the structure you never touched as where most of the remaining gain sits. I came back to that plant from a different angle in the Team Member Handbook post.
Are you doing Agile, or have you become Agile?
The difference seems pedantic at first…
You are doing Agile when you’ve changed your tools and processes. This is relatively easy to do but doesn’t offer much in the way of benefits. You’ve becomeAgile when you’ve changed you structure and culture too. This is relatively hard to do, but offers significant benefits.
Agile isn’t just a process. It’s a complete framework that brings together a shift in culture, structure, and processes. This framework is supported by tools such as Rally and other Agile Lifecycle Management (ALM) tools.
[May 2026]The three-part frame (delight, value, good) still anchors how I think about products and AI tools alike. The “create value” section has matured into Reinertsen’s flow economics and cost-of-delay weighting; the “do good in the world” section has become more important, not less, as AI raises the externalities question.
My elevator speech back then went like this: “I design, develop, and ship innovative products that delight customers, create value, and do good in the world.”
Those last three components—delight customers, create value, and do good in the world—are the three most important aspects of a successful product. Here’s why I think so.
[May 2026]Silent Lens didn’t ship; for off-grid mesh comms today I’d point at Meshtastic, with the caveat that Meshtastic isn’t hardened against adversarial regimes the way this design aimed to be. The public-key crypto for evidence, immune-system trust gradient, and courier-fallback delivery were all aimed at a much harder problem, and the architecture still seems plausible to me.
I attended the Google I/O Extended “Develop for Good” hackathon in San Francisco in late June. We were asked to create a solution for one of three challenges:
Google Politics & Elections: Citizen Engagement for Politics & Elections
Google Ideas: Conflict Reporting for Blackout Situations in Repressive Regimes
[May 2026]This argument still holds. Both “ship quickly and often” and “defer commitment” remain true today; cost-of-delay weighting (Reinertsen, Fox & Gregory) and real-options framing are the sharper economic vocabulary that grounds them.
[August 2026]“They build slowly and test often” is the practitioner’s version of an argument I make more formally in Chapter 7 of Come Prepared to Die: a plan held as a commitment kills the learning, and a plan held as a hypothesis invites it. The marshmallow is what a premature commitment feels like when it lands.
Spaghetti and Twine
Many of you will be familiar with Peter Skillman’s Marshmallow Challenge, an exercise frequently given to teams and business school students. Teams of four are given 20 pieces of spaghetti, 1 yard of tape, one yard of twine, and a marshmallow. They are then given 18 minutes to build a free-standing structure that places the marshmallow as high off of the table as possible. The team with the highest marshmallow wins.
If you haven’t seen it already, Tom Wujec’s TED talk is a good place to learn about the challenge. And if you haven’t introduced your team(s) to it, take 45 minutes out of one of your days to administer the challenge and see what revelations you get.
[May 2026]I still think rigid priority-number columns mislead, but I no longer believe the fix is sitting down with stakeholders to sort the list by hand. Today I’d weight by cost of delay (Reinertsen’s WSJF, Fox and Gregory’s economic framing) and treat ordering as a live conversation grounded in numbers, not as a one-time stack-rank artefact. The instinct toward relative priority was right. The toolkit was thin.
[August 2026]Cost of delay is the variable Chapter 4 of Come Prepared to Die names, and this 2008 post was reaching for it without the word. What I argue here is only that relative order beats absolute priority. The chapter adds the economics that argument was missing: in high-variation work, keeping everyone busy makes the system slower, and queues are what you are managing whether you know it or not. The Customer is the Marshmallow makes the companion argument: defer the commitment and test before you build on top of it.
More specifically, I hate numbers or letter representations of priorities when it comes to product backlogs.
It’s a common strategy, even in Scrum. (Henrik Kniberg’s wonderful scrum book talks about a product backlog where higher priority items get higher priority numbers, preventing the “if this is critical and priority 0, what is ultra-critical? priority -1” issue.)
So why the hate? Simple – they do a lousy job of actually priortizing tasks. How many times have you encountered a product backlog where there were several items that were all of critical importance? How is this truly helpful?
Think of it in this way – what if half the items in your email inbox were of CRITICAL priority? At this point, what value does this tag add? At the end of the day, you’ll have to choose ONE thing to do next. What will it be?
I therefore argue that it’s exactly this hard decision that needs to be made earlier in the process, with the stakeholders who will wonder why this critical priority issue took precedence over that critical priority issue.
The real issue is that priority values attempt to apply a rigid metric of ABSOLUTE priority when the only thing that matters in the real world is RELATIVE priority – what do we do next? Even if you have the ability to complete work in parallel (e.g., more than one developer), you still need to figure out what those n people will do next.
Therefore, I propose that we kill the concept of priority values in the agile workplace.
Take your product backlog, remove the priority column, and sit down with the stakeholders. Don’t walk out of the meeting room until every item is sorted in order of relative priority.
The rest is easy: in your next sprint planning meeting, figure out how many story points you have available and work down from the top of the list. There are only two exceptions:
When the developers believe that two pieces of work are similar enough to realize greater efficiencies if completed together. If this happens often, you need greater developer involvement in the priority setting meeting.
When the remaining story points don’t support the next priority item. For instance, suppose there are 3 remaining story points but the next item in the product backlog requires 5. It’s OK to scan down a little and take the next item at or below three points.