You shipped more this quarter than last. Your customers are no better off. Which number is the work actually being judged on?
By now you know the move: find the axis where your claimed pole and your lived pole pull apart, then give yourself to that analysis. If output-versus-outcome is already the tension eating your quarter, this is your chapter. If it isn’t, read it lightly and save your energy for the axis that is.
Subordinating non-constraints was chapter 8’s move. It pays off only when the constraint serves the right number. Output and outcome are different numbers, and the gap between them is where feature factories live.
🎧 Prefer to listen? This chapter is narrated in my own voice with ElevenLabs on Spotify (~18 minutes).
The two poles
Pole A: output. The unit of management is velocity, units shipped, features released, work completed.
Pole B: outcome. The unit of management is customer progress, market position, mission impact.
This axis is well known among engineering leaders. The vocabulary—feature factory, output vs outcome, customer progress—has been in circulation for well over a decade.
Where Pole A is right
Pole A is right when output and outcome have been shown to track each other tightly in the relevant domain. In a sales organisation with a mature product, deals booked may convert predictably to revenue over months. In a fulfilment operation, units shipped may equal customer demand met, provided return rates remain low and margins stable. In a regulatory pipeline, filings submitted may track approvals secured when the regulator’s response time is well-characterised over multiple cycles.
Most senior leaders of software-dependent organisations came up through some version of Pole A. The metrics were ideal in their context. Deals booked, units shipped, and features released were proxies that tracked the customer outcome closely enough that managing the proxy amounted to managing the outcome.
None of that discipline was wasted. Counting output is part of what makes anything ship at all. A team that can’t count what it has produced can’t tell learning from activity.
Pole B does not abandon counting. It asks you to notice when output stops tracking the outcome. Output was a clean proxy until the conditions that made it clean stopped holding; after that, counting output means counting the wrong thing.
Where Pole B is right
Pole B is right whenever the firm produces more output and gets worse outcomes. This is the empirical pattern in feature-factory product organisations. The velocity’s high and the customers still leave. A roughly similar pattern appears across SaaS organisations under aggressive growth pressure, regulated product organisations whose feature roadmap is committed years in advance, and platform organisations whose internal customers are the engineers themselves.
Clayton Christensen’s Jobs to Be Done theory, laid out in Competing Against Luck, gives the central reframe. The customer isn’t buying your feature. The customer is hiring your product to make progress on a job.
The classic example is the milkshake a fast-food chain sold in surprising numbers before 9 a.m. to commuters buying nothing else. Commuters hired the milkshake to keep a long, boring drive interesting, one-handed, and to hold off the mid-morning hunger. The milkshake was hired because it outlasted every competitor—a thick shake takes twenty minutes through a thin straw—and the cup fit the cup-holder.
Improving the milkshake meant making it last longer, not making it taste better. The Pole A move (survey the customers, ask what flavour they want) produced flavour iterations that didn’t move sales. The Pole B move (understand the job) produced an actionable insight.
Marquet: red work and blue work
David Marquet’s Leadership Is Language names an operational discipline that fits the output-outcome distinction. Red work is execution under pressure: doing what has been decided. Blue work is judgement and pause: deciding what to do.
The Industrial-Age inheritance divided planning from doing. A few people decided; everyone else executed. Marquet recasts those activities as blue work and red work, then argues that work now requires everyone to do both.
He also replaces the reactive language of convincing, coercing, complying and conforming with stated intent and commitment to action. Proving and performing give way to improving and learning, while leaders exchange invulnerability and certainty for vulnerability and curiosity.
The Pole A leader who manages to features-shipped runs in pure red mode. The team decomposes the goal into tasks, executes them, and reports completion. Blue work has no place to happen because the language of the work has no room for judgement.
The Pole B leader inserts blue work into the cadence. The team pauses, checks whether the output delivered the outcome it was meant to serve, and changes the next round of red work based on what it learned.
We committed to ship X by Q3 closes the loop on red work. We hypothesised that shipping X would produce outcome Y by Q3; what did we learn? opens the loop into blue work.
Same artefact. Different relationship to the work and leader.
Commander’s intent: name the outcome, not the steps
Marquet translated an older doctrine into a protocol that leaders can use in everyday decisions. His specific articulation in Turn the Ship Around!, the “I intend to…” protocol on the USS Santa Fe, is the most accessible contemporary form.
The mechanism is small and specific. The subordinate replaces “request permission to…” with “I intend to…”, states the intent, gives their rationale, invites veto, and proceeds unless the leader vetoes the decision. The protocol shifts the cognitive load from the leader to the person closest to the work.
Name the outcome, not the steps. A plan that specifies the steps dies the moment events outrun it; the statement that survives is the goal and desired end-state, compact enough that people down the line can improvise toward it.
Innovation accounting
Eric Ries hands the output-counting leader an accounting discipline for Pole B. Vanity metrics rise regardless of whether anything is working; actionable metrics change what the leader does next.
Here’s the test I use to tell them apart. Pick the metric. Predict what it will do over the next two weeks or months. Look at it. Did the prediction match? If it didn’t, what would you do differently?
If the answer to that last question is nothing, the metric is vanity. A Pole A leader running on vanity metrics isn’t running on data. They are running on the appearance of data.
Innovation accounting, Ries’s term for tracking learning per unit of investment, converts Pole B into language the Pole A board can use without losing accountability. It is the chapter-7 board script in a different currency: commit to learning a named thing, by a date, at a cost, then report what was learned and what to test next.
Cagan: valuable, usable, feasible, viable
Marty Cagan’s outcome test is the simplest I have seen for product work. A feature is worth shipping if it is valuable, usable, feasible, and viable.
Valuable means it produces an outcome the customer cares about. Usable means the customer can access that value. Feasible means the team can build and operate it within the firm’s constraints. Viable means the business itself can sustain it.
The failure modes are easy to recognise. A Pole A team may optimise for feasible, building what it can build and shipping features that are technically clean but irrelevant to customers. The most user-research-heavy teams sometimes over-rotate on valuable and ship features the team can’t operate at scale.
A feature is worth shipping only if it lands all four. Most feature-factory output meets only one or two.
Edmondson’s intelligent failure
Organisations can absorb output vs outcome as a slogan and keep measuring the same old work. Fail fast becomes operational only when a failure meets Edmondson’s four criteria for an intelligent failure, which chapter 10 walks in full.
Intelligent failures aren’t errors. There was no “right” way to do the work in the first place, a point Amy Edmondson presses. A Pole B leader managing to outcomes runs at a higher cadence of intelligent failure than a Pole A leader because learning has a cost.
A Pole A leader can adopt the fail fast slogan while keeping the old output metrics. In that system, teams produce the appearance of intelligent failure and count the failures rather than what they learned from them.
A representative case
Picture a SaaS product team that has been measuring features shipped per quarter for two years. The number goes up steadily. The customer churn rate also goes up steadily.
The CEO can’t reconcile the two numbers. Engineering leadership points at the customer success function; the customer success function points at product. The board is being told we are shipping faster than ever in the same quarter that renewal at 18 months has dropped twelve points.
The Pole B move arrives through a customer advisory board. Several customers on the board, on separate calls, name the same pattern: the product has grown surface area faster than they can absorb it. The features they asked for two years ago are buried under features they didn’t ask for, and the team they originally trained on the platform has churned.
Less would be more useful than more.
The retrospective surfaces the structural problem. Every feature shipped had been valuable, usable, feasible, and viable at the time it was scoped. Nobody had re-evaluated it against the integrated product experience.
The features exist, and customers sometimes use them. The outcome is being eaten by their accumulation. The team doesn’t lack discipline; it has been managing the wrong unit for the conditions that developed.
The fix wasn’t to abandon shipping. The team introduced a quarterly outcome review alongside the existing roadmap review and asked, for each major feature shipped over the past four quarters, whether customers were using it, whether that use produced the outcome the feature was supposed to deliver, and whether the maintenance cost still tracked the value.
Half the features failed the test. A third were quietly deprecated. The next quarter’s roadmap was leaner by a factor of two, and the customer cohort that had been churning stopped churning within six months.
The engineering leader running the original output count wasn’t a bad leader. The previous CEO had measured them on output. The investor decks had measured them on output. The internal performance reviews had measured them on output. The system around the leader had reinforced Pole A for years.
So name what the discipline got the team—predictable cadence, investor confidence, a high-trust engineering culture—and then name what the new conditions required. The leader didn’t need a different character. The leader needed a different measurement system, and time to adjust to it.
The diagnostic move
Three questions for last Tuesday’s launch celebration.
- Which pole was I claiming? In the language I used about the launch, did I name what the customer would now be able to do, or did I name what the team had shipped?
- Which pole did the team’s behaviour show? Look at where the team’s attention went in the week after the launch. Did it go to customer use or to the next item on the roadmap?
- Which pole does the work actually require? If output and outcome track tightly in your domain, Pole A is legitimate. If you are shipping more and the customer outcomes are flat or worse, Pole A has stopped being a proxy and started being a substitute.
The board deck locks in this distinction, and that deck belongs to your CEO. If you run engineering, the metric your CEO uses to describe your function to the board governs the output-versus-outcome question.
Put the metric that currently judges your function beside the metric that tracks customer progress. In most software-dependent organisations, these are different numbers. A CEO who describes engineering to the board in features-shipped vocabulary creates Pole A incentives three layers down, usually without intending to.
Ask for one sentence in the next board update that names what customers can now do, not what the team shipped. That is the Pole B move at board level. The team will read the update.
Give your CEO this sentence: “Customers are now doing something they could not last quarter; the features we shipped are how we got there.”
The exercise
Rewrite a recent feature as a job to be done. Christensen’s test is specific: a job is the progress a person is trying to make in a particular circumstance; state it in verbs and nouns, and check that candidates for the job come from different product categories.
The job-story template—“when [situation], I want to [motivation], so I can [expected outcome]”—is a usable scaffold. It comes from the job-story tradition, not from Christensen. The format forces the team to name what the customer is hiring the feature for.
Notice how often the feature solves the firm’s job rather than the customer’s. When my quarterly performance review is approaching, I want to ship a visible feature, so I can demonstrate productivity to my manager is a job. It belongs to the employee, not the customer.
Run a companion vanity-vs-actionable metric audit. List your top five metrics. Mark each as vanity (it goes up regardless of whether anything is working) or actionable (it changes behaviour when you look at it). Defend each judgement to the team.
The exercise usually finds that three of the five are vanity. The conversation about which three is the work.
If leader-follower dynamics turn out to be your live axis, run Marquet’s “I intend to…” protocol for a week. Chapter 17 lays out the full version. Even one week of replacing “request permission to…” and “recommend that we…” with “I intend to…” shifts who carries the cognitive load of the decision, and the effect is visible in days.
Going upstream
In-text: Clayton Christensen, Competing Against Luck, on Jobs to Be Done; David Marquet, Turn the Ship Around! and Leadership Is Language, on red/blue work and the language of intent.
Also touched: Eric Ries, The Lean Startup, on innovation accounting and vanity metrics; Douglas Hubbard, How to Measure Anything, on the test that a measurement only earns its keep if it changes a decision; Marty Cagan, Inspired, on valuable / usable / feasible / viable; Jeff Patton, User Story Mapping, and the talk Output vs Outcome & Impact (11 min, watch); Amy Edmondson, Right Kind of Wrong, on the four criteria for intelligent failure (developed in chapter 10).
Go deeper: Chip and Dan Heath, Made to Stick, on commander’s intent in organisational-messaging terms; USMC, MCDP-1 Warfighting, Ch 4, on mission tactics and commander’s intent, with the German antecedent Auftragstaktik from nineteenth-century Prussian military reform; Frederick Taylor, The Principles of Scientific Management (1911), on the planning/doing division Marquet recasts as blue and red work; Alan Klement, Replacing the User Story with the Job Story (2013), on the job-story form, which originated at Intercom.
Manage to outcomes and some of them will miss. Watch what you defend when yours does. The call you made was a testable bet; the thing you protect in the review is usually your identity, and the second job quietly eats the first. Chapter 10 sits with that review, the one after a decision fails, and asks what the miss costs you next time: whether you spend it shielding the person who made the call or finding out what the call was worth.
Leave A Comment