The Instrument

Six chapters handed this one an unpaid bill.

Chapter 5 needed a test for telling a complicated situation from a complex one and offered three heuristics with a note saying they were provisional. Chapter 11 needed a filter that could be held accountable and said the clinical literature is fifty years of very capable people failing at it. Chapter 14 needed a way to tell automation that serves its owner from automation that serves somebody else. Chapter 15 needed the inventory step that converts a permission into a procedure. The chapter on internal contradictions needed a maintenance budget with a number on it. Chapter 17 handed over two rules the creed did not previously have.

What follows is one instrument that takes all of it. Seven gates, in order, each with a question that has an answer rather than a mood.

Two things about what it is not, because an instrument oversold is worse than none.

It does not tell you whether a formula will work. It tells you whether to try, and what the trying will cost, and what has to be true for the result to be trustworthy. Those are answerable in advance. Whether the thing works is answerable only afterward.

And it is deliberately slow at the front. Most of its value is in the first three gates, which stop projects, and stopping a project is the cheapest good outcome available and the one nobody gets credit for.

* * *

Gate one. Does this decision actually repeat?

The question sounds trivial and it fails more candidates than any other gate.

A great deal of work that looks repeated is a sequence of one-offs sharing a name. Six installations at six sites are six jobs wearing one word. Twelve contract reviews are twelve contracts. The name is the same and the decision is not, and a rule built across them will be a rule built across a category that exists only in the vocabulary.

The test is not whether the label recurs. It is whether the inputs recur: whether the thing you would be reading to make the decision takes similar values across instances, in a way that lets you say what the typical case looks like. If you cannot describe the typical case, there is no typical case, and the formula has nothing to be about.

Two numbers are worth having before going further, and both are usually available in ten minutes. How many times a month does this happen? How long does it take each time, including the deciding rather than just the doing?

Below roughly a dozen instances a month, almost nothing clears the economics, because a rule has to be specified, validated, maintained, and audited forever, and the person deciding is already employed. Chapter 15's meta-analysis found nearly half of comparisons ending in a tie between a person and a formula, and a tie is a win for leaving it alone. The cases that pass this gate are the high-volume, low-drama, invisible ones that nobody complains about, which is why they are also the ones nobody proposes automating.

Failing here is the most common outcome and the correct answer is to stop.

* * *

Gate two. Which domain is this in?

This is the gap Chapter 5 named as the creed's most expensive, and it is expensive because from inside, a complex situation looks exactly like a complicated one that nobody has analyzed hard enough yet.

Three questions, none of which requires understanding the subject matter, which is what makes them usable by somebody who is not the expert.

Does anyone predict this in advance and turn out to be right? Not explain afterward, which is always available and always convincing. Say beforehand and be correct more often than chance. Where such people exist, cause and effect are knowable and you are in the ordered domains whatever the situation feels like. Where the most expert-seeming people are only ever right in retrospect, that is the signature of complexity, and it is detectable without any domain knowledge at all.

Does the same intervention produce the same result twice? Run it, wait, run it again. In the ordered domains the second run resembles the first. In the complex domain the system has responded to the first intervention and the second is acting on a different system.

Can an experienced person explain it well enough that a newcomer handles it correctly? Complicated problems transfer through explanation; that is what makes expertise teachable. Complex ones do not, and the tell is the experienced person who cannot say why they did what they did, because there was no why in existence before the doing.

Two of three pointing toward complex is enough to stop. Not because nothing can be done, but because what should be done is not this creed: small safe experiments, variation preserved rather than suppressed, patterns allowed to show themselves. Applying the method here does not merely underperform. It removes the variation that was the only available instrument.

One warning on this gate. It is the one people will fail dishonestly, in both directions. Somebody who wants the project will find the situation complicated. Somebody who does not will find it complex. The three questions help because they are about observable facts rather than about the situation's feel, and the first one in particular is checkable by asking around.

* * *

Gate three. Are the conditions observable, stable, and not manufacturable?

These are the three the formula chapter arrived at, and they are the technical core.

Observable means a person can look and answer without interpreting. Not is the customer frustrated, which requires a judgment that varies by judge. Has the customer called more than twice about this, which does not. This gate is where most badly designed formulas die, and the death is usually delayed, because a condition that requires interpretation will produce answers immediately and they will disagree with each other invisibly for months.

Stable means the mapping from condition to action holds still at least as long as the rule will run before somebody looks at it again. This is a claim about a rate, and the honest way to answer it is to name the last time the mapping changed and how anybody found out.

Not manufacturable is Goodhart's condition, and it is the one nobody states. Ask who benefits from the condition being true, and whether they can bring it about. If they can and they do, this is not a formula. It is an incentive wearing a formula's clothes, and the observable regularity you built it on will collapse precisely because you built it on that regularity.

The manufacturability test has a second form that catches more cases: ask what the laziest available way to satisfy this rule is, and whether it resembles the thing you actually want. If a status stops a clock, parts get ordered. If contact means an email was sent, emails get sent.

Failing observable is usually fixable by going down a level. Failing stable means building a rule with a review date attached. Failing manufacturable means either changing what the rule reads or accepting that the rule and the measurement have to be designed by people who are looking at each other.

* * *

Gate four. What does the person know that the rule will not?

This is Chapter 15's inventory, and it is the gate that most changes what gets built.

Ask whoever currently makes the decision what they would need to know in order to make it, and write the answers down item by item, in specifics. Not I know these customers, which is unfalsifiable. This site manager will not accept that crew, which is a fact with a truth value.

Then sort the list into three piles.

Items that could be captured and simply are not. This is almost always the largest pile, and it is the discovery that matters, because the correct response to the human knows things the system does not is usually to put those things in the system. Site restrictions are data. Duration corrections are data. Every item moved out of this pile makes the eventual rule better and makes the person's remaining contribution clearer.

Items that could be captured at a cost nobody will pay. These are real and they are a design decision rather than an impossibility. Chapter 3's conclusion applies: go deep where the facts are stable and verifiable, stay deliberately shallow where they move faster than anyone will maintain them, and make the shallowness explicit so the field says ask instead of quietly reporting last quarter's number.

Items that cannot be captured. This is the genuine residue and it is smaller than the first list every time I have run this. What remains here determines the shape of the rule: the rule proposes and the person disposes, with overrides recorded, because an override rate is the only instrument that measures how wrong a rule is from inside the system running it.

If the residue is large and irreducible, this gate fails and the decision stays manual, and it stays manual for a stated reason that somebody else can evaluate.

* * *

Gate five. What does the current arrangement do for the people doing it?

This gate did not exist until Chapter 17 and it is the one I most regret not having had for the last decade.

The question is not whether people will accept the change. It is what the existing way of working is producing besides its output. The Durham longwall was technically superior and it destroyed a social structure that had been carrying load nobody had put in the drawing, and the output never arrived.

Three things to look for, each of which has shown up somewhere in this book as an unmeasured loss.

What does the current work make somebody look at? A classification that requires reading a request whole produces a picture of the account as a byproduct. Automate the classification and the picture becomes optional, and optional work stops happening, and nobody decides to stop.

What does it make somebody practice? A skill exercised as part of the ordinary flow is a skill available on the bad day. Remove the exercise and the skill is demanded at the worst possible moment after months of disuse.

What does it make people do together? Work that requires two people to talk produces a relationship that carries other traffic. Splitting it across a system and two shifts removes the conversation, and the traffic it was carrying does not reroute, it stops.

None of those is a reason not to proceed. All three are line items that belong in the design and in the cost, and the practical form is the same in each case: if the automation removes the reason for a thing that was worth having, put the thing back deliberately, on a schedule, as part of the build rather than as a remedy afterward.

The failure mode this gate prevents is the one where every measure improves and the organization gets worse, which is not rare and is nearly impossible to diagnose after the fact.

* * *

Gate six. Can the person the rule governs read it and reverse it?

The second rule out of Chapter 17, and the shortest gate.

Can the person subject to this rule find out what it says, in terms they can evaluate? Not the source code. The rule: you get assigned jobs this way because of these four things, written where they can see it.

Can they override it, at a cost low enough that they will, with the override recorded?

Can somebody turn it off, today, if it turns out to be wrong?

Three yeses and the thing is a tool. A no on the first means the people governed by it cannot tell you it is wrong, which means the override instrument from gate four does not exist, which means the maintenance instrument from gate seven has nothing to read. A no on the third means it is not a formula anybody installed; it is a condition they live in, and Scott's analysis applies at whatever scale it operates.

This gate has no partial credit and it is cheap to pass. Its only enemy is that writing down what a rule does invites people to argue with it. That is what the gate is for, and it is experienced as a cost.

* * *

Gate seven. What is the maintenance allocation, and who is paying it?

Every rule that runs unattended decays, and Chapter 16 established that a decayed rule and a working rule are behaviorally identical from inside, because conformance tests compare behavior to the rule and the rule is being followed perfectly.

So the last gate is a number. How much human attention, per month, permanently, is allocated to comparing this rule against the world? Not to improving it. To finding out whether it is still true.

Four instruments, and a design should carry at least two.

Expiry on any underlying fact that carries a date. Cheapest, runs unattended, covers little.

Override rate, which requires that gate six passed and that overrides are recorded rather than merely possible. A rising rate on one rule is that rule reporting its own obsolescence.

The case that will not fit, captured rather than discarded. Somebody forcing a request into the nearest category is the earliest available signal and it currently lives for four seconds in one person's head. Making it recordable costs one field.

The manual sample. Pull a handful of the rule's outputs a month and check them against the world by hand. This is the only instrument with no blind spot in principle and it is the one every efficiency argument in this book is aimed at eliminating.

The allocation is small. Twenty checks a month is not a meaningful fraction of what a rule running thousands of times saves. The reason it does not survive contact with a budget is that it looks like waste on every metric an automation project is evaluated on, and its entire purpose is to discover that the project was wrong.

If this allocation is not funded, the honest description is that the organization has chosen to run the rule until it fails visibly, and that choice should be made out loud rather than by omission.

* * *

Run it once, on a real case, so the gates have something to bite.

The case is crew scheduling. A queue of jobs needs assigning: which crew, which day, in what order. There is distance, duration, skill, availability, priority and travel time, and an optimizer can beat a person on every one of those in about a second. It is the most automatable-looking problem in the building.

Gate one. Forty to sixty assignments a week, every week, each taking a few minutes of real deliberation. The inputs recur and the typical case is describable. Passes easily, which is the first thing that makes this a serious candidate rather than an enthusiasm.

Gate two. Does anybody predict this in advance and turn out right? Yes: the dispatcher says on Monday which jobs will slip and is correct more often than chance, and everybody in the building knows it. Does the same intervention repeat? Yes: assigning a crew to a site produces roughly the same result on Tuesday as on Thursday. Can an experienced person explain it to a newcomer? Mostly, with a residue. Three for three toward the ordered domains. Passes.

Gate three. Observable: distance, duration estimates, certifications, and stated availability are all lookups. Passes. Stable: the mapping changes slowly, as crews are hired and certifications expire, which is a rate somebody can name. Passes with a review date attached. Not manufacturable: this is where it gets interesting, and the answer is partly. If the rule reads a crew's stated availability, and a crew learns that stating less availability means easier weeks, the condition is manufacturable by the people the rule governs. The fix is not a better definition. It is that availability comes from a schedule of record rather than from the crew's own assertion, which is a design change that gate three forced and that nobody would have thought of while admiring the optimizer.

Gate four. The inventory, run with the dispatcher. Four items came out. One site manager will not accept a particular crew. A job listed as one day will take two, because she talked to the site that morning. One technician needs short days this week for something at home. And one customer has been patient twice and will not be a third time.

Sorted: the site restriction is capturable and became a field. The duration correction is capturable and became a field. The technician's week is capturable in principle and should not be, because it would mean recording somebody's family situation in a work system in order to save a few minutes of scheduling. The customer's patience is the genuine residue: real, decision-relevant, and not a thing anyone should try to make into data.

So the rule proposes and the dispatcher disposes, overrides recorded. That is a different design from the one the optimizer's demo implied, and gate four is the entire reason.

Gate five. What does the current arrangement do for her besides producing a schedule? It makes her look at every job on the queue each morning, which is how she knows which ones will slip, which is the prediction that passed gate two. Automate the assignment and she stops reading the queue, and the capability that made her useful at gate two disappears about four months later.

That is not a reason to stop. It is a line item: the morning review stays, as a deliberate step, after the optimizer has proposed. It costs her fifteen minutes and it preserves the thing the whole design depends on.

Gate six. Can a crew find out why they got the jobs they got? Before this exercise, no. Now, yes, in four sentences on the assignment itself. Can they push back, cheaply, with the pushback recorded? Yes, and that record is the override instrument. Can somebody turn the optimizer off today? Yes, because the manual path was kept rather than decommissioned, which cost something and is the reason the answer is yes.

Gate seven. Twenty assignments a month get checked by hand against what actually happened. Override rate gets watched per field rather than in aggregate, because an aggregate rate hides which part of the rule has gone stale. The forced-fit case gets a field. Fifteen minutes a week, named in the build, funded.

The result is not the system the optimizer demo promised. It is slower to build, it keeps a person in the loop on purpose, and it has an ongoing cost line. It is also the only version I would trust in three years.

* * *

Three things this instrument cannot do.

It cannot tell you whether the formula will be any good. Every gate is about whether a formula is appropriate, and appropriateness does not imply quality. A rule can pass all seven gates and encode a bad policy perfectly.

It cannot detect a definition that is deep, well-maintained, consistently applied, and pointed at the wrong quantity. That failure passes every test in this book, because every test in this book compares a thing to its own specification.

And it cannot survive being turned into a form. The moment these seven gates become a checklist with boxes, gate two becomes a box somebody ticks, and the three questions underneath it stop being asked. That is this book's own machinery turned on this book, and by the argument of Chapter 8 I should expect it: the moment a gate governs whether a project proceeds, the gate stops being a neutral observation and becomes a thing people produce.

I do not have a defense against that beyond saying it out loud. The instrument works as a conversation and degrades into theater as a document, and the difference is whether anybody in the room is empowered to answer gate one with no.

* * *

A note on what happens when a gate fails late.

The gates are ordered so that most failures are cheap, and some failures cannot be discovered until the thing is running. A condition turns out to be manufacturable only after somebody manufactures it. A social cost at gate five shows up as a capability that quietly stopped, eight months later, in a person who cannot tell you what they lost.

For those, the instrument has one thing to offer and it is gate seven doing double duty. The maintenance allocation is not only a decay check. It is the mechanism by which a late gate failure gets discovered at all, because the person sampling outputs against the world is the only person positioned to catch a rule that has started rewarding something nobody intended.

Which means the sampling should not be a spot check on accuracy alone. The person doing it should be asked two questions rather than one. Did the rule produce the right answer, and is anyone behaving differently because the rule exists. The second question is the one that catches gate-three failures after the fact, and it costs nothing to add, and in three years of doing this I have never seen it on anyone's audit sheet including my own.

* * *

Two failure modes of the instrument itself, both of which I have produced.

The first is running the gates in the wrong order. They are sequenced deliberately: the cheap disqualifying questions come before the expensive investigative ones, so that a project dies at gate one for eleven minutes of work rather than at gate four after a fortnight of interviews. The temptation is always to start at gate four, because gate four is the interesting one and involves talking to people who know things. Starting there means the inventory gets run on a decision that does not repeat, and everybody's time is spent, and the finding is unusable.

The second is treating a gate failure as an obstacle to be argued past. A failure is information and it is the most valuable output the instrument produces, because the alternative is finding out after the build. Gate three telling you a condition is manufacturable has just saved a year of a rule quietly training people to game it. Gate five telling you the current work makes somebody look at something has just told you what the automation will cost in a currency nobody was tracking. Neither is an objection to be overcome. Both are the design.

There is a version of this instrument that gets used as a defense, where a person who does not want to change anything runs the gates and finds a failure at every one. That is the boundary ism's structural attraction arriving in procedural form, and the only guard I know is that a gate failure has to name a specific fact rather than a concern: not the conditions might not be stable but the mapping changed twice last year and here is when. A failure that cannot name a fact is a preference, and preferences are legitimate and belong in a different conversation, which is exactly the distinction Chapter 15 ended on.

* * *

One shape worth noticing across all seven.

Gates one through three ask whether the world will hold still enough for a rule. Gate four asks what the rule would be missing. Gates five and six ask what the rule does to people. Gate seven asks who pays to keep it true.

Only the first three are about the technical problem, and those are the ones every version of this creed had before I started writing it down. The last four came out of the objections, which means the instrument that took me eleven chapters to build is more than half made of things the creed was wrong about.

That is the most honest summary of this book I can give. The method was right about the machinery and silent about everything the machinery touches, and the silence is where all the damage was.