"Where a data-based decision will not produce a genuinely better result, leave it manual."

Where a data-based decision will not produce a genuinely better result, leave it manual.

The weekend list is six items and it is fully specified. Laundry. Floors vacuumed and washed. Grass cut and edged. Bathroom. Dishwasher run and emptied. Groceries every couple of weeks.

What is not specified is anything about when. Not the order, not the day, not which of them happens before the coffee and which gets pushed to Sunday afternoon. That is deliberate, I have been doing it deliberately for years, and it is the single clearest instance in my life of the rule this chapter is about. I could schedule those six things. I have the tools. I have built considerably more complicated schedules for other people's work. I do not schedule my own weekend, because a scheduled weekend is not a better weekend, and the hours the list does not claim are worth more to me unclaimed.

That is the rule working, and it is also the version of it that cannot defend itself.

Because here is the problem I have avoided for as long as I have had the sentence. A person who has never automated anything in their life, who resists every system anyone proposes, who calls their avoidance intuition and their disorganization freedom, would describe their behavior in exactly the words I just used. A data-based decision would not produce a genuinely better result here. Same sentence. Completely different thing happening underneath, and nothing in my formulation can tell them apart.

A boundary that admits every case that wants admission is not a boundary. It is a permission slip with a technical vocabulary, and I have watched it issued, and I have almost certainly issued it to myself.

* * *

Stated as a rule:

The variables are a decision, the quality of the result it produces when made by a person, and the quality it produces when made by a rule operating on data.

The quantity compared is the difference between those two, and where a data-based decision will not produce a genuinely better result, leave it manual says: where the difference does not favor the rule by enough to be worth the cost of building and maintaining it, do not build it.

The failure signal, in the direction I usually worry about, is a formula producing answers nobody trusts and everyone works around, which means the conversion happened where it should not have.

The failure signal in the other direction is the one this chapter has to add: a decision made by judgment, repeatedly, that nobody can explain, that comes out differently depending on who made it, and that nobody has ever tested against a rule. That is the smell of a decision that was left manual by default rather than by finding.

The domain of validity is wherever the comparison has actually been made. The sentence is a conclusion, and I have been using it as a premise.

The load-bearing phrase is genuinely better, and the word genuinely is doing what almost does in the formula ism and unnecessarily does in the vagueness ism. It is where I stopped examining. Better on what dimension, measured how, over what horizon, compared against what baseline? The sentence does not say, which is what lets the same words serve both me and the person who has automated nothing.

* * *

There is a serious intellectual tradition behind leaving things alone, and it is worth laying out before the evidence takes it apart.

James Scott's account is the broadest. His subject is what happens when a central authority needs to see something in order to manage it, and what gets destroyed in the making-visible. The premodern state, he writes, was in many crucial respects partially blind: it knew precious little about its subjects. Permanent surnames, standardized weights and measures, cadastral surveys and population registers were the instruments that fixed that, and every one of them is a standardization of variables in the sense this book has been advocating for four chapters.

His scientific forestry case is the one that should worry anyone who builds systems. German foresters achieved genuine rigor by simplification: the actual tree, with its vast number of possible uses, was replaced by an abstract tree representing a volume of lumber or firewood. That abstraction made the forest calculable, which made it manageable, which produced real gains. It also produced ecological collapse a century later, because everything the abstraction dropped turned out to be doing something.

Scott's term for what gets dropped is metis: practical knowledge built by experience, particular to its setting, adaptive to circumstances that shift. He contrasts it with techne, which is codifiable and universal, and his illustration is a cook being told to heat the oil until it is almost smoking. That instruction is useless to someone who has not stood over oil, and complete to someone who has. Ten times ten is a hundred everywhere and forever. Almost smoking is not.

And he names four things that combine to produce catastrophe: administrative ordering of nature and society, a high-modernist confidence in technical progress, authoritarian power willing to impose the design, and a civil society too weak to resist. The first two are the ones available to anyone building systems inside an organization. The second two are the ones that make it lethal, and their absence is why most bad system design is merely wasteful.

Gary Klein's fireground commanders, who appear in the definition chapter, are metis with a stopwatch on it: twenty-six commanders averaging twenty-three years, a hundred and fifty-six decision points, fewer than twelve percent showing any comparison of options at all. Dave Snowden and Mary Boone's complex domain says that where cause and effect cohere only in retrospect, imposing order fails. Ashby's requisite variety says a rule with fewer distinct states than the world produces will answer some situations with an action built for a different one.

All of that supports the rule. It is a coherent, serious, well-populated tradition, and the next section is where it meets the thing that has been sitting in the psychological literature since 1954.

* * *

Paul Meehl published a short book that year comparing two ways of making a prediction about a person. One is clinical: an expert takes in the information and forms a judgment. The other is mechanical: the information goes into a formula, often a very simple one, and the formula produces the answer. He reviewed the studies available and found the formula winning or tying nearly everywhere.

That finding has been replicated and extended for seventy years and the definitive synthesis is Grove, Zald, Lebow, Snitz and Nelson's meta-analysis published in 2000, covering a hundred and thirty-six studies.

Their headline: mechanical prediction was about ten percent more accurate than clinical prediction on average, with a weighted mean effect size of point zero eight six.

Forty-seven percent of studies showed substantial superiority for mechanical prediction. Forty-eight percent showed the two approximately equal. Six percent, which is eight studies out of a hundred and thirty-six, showed substantial superiority for clinical judgment.

Eight out of a hundred and thirty-six. That is the empirical base rate for where a data-based decision will not produce a genuinely better result, leave it manual being right in the settings where somebody bothered to run the comparison, and it is small.

Two honest caveats in the rule's favor, and then one finding that destroys the easy version of it.

The average effect is modest. Ten percent is not annihilation, and nearly half the studies are ties, which means that in about half of these comparisons a person performs as well as the formula. If the formula costs something to build and maintain and the person is already there, a tie is a win for leaving it manual, and the rule survives in every one of those cases on cost grounds rather than accuracy grounds.

And the domain is specific. These are predictions about human beings made by clinicians and other experts, in settings where an outcome could be scored. That is not every decision. Whether it generalizes to deciding how to spend a Saturday is an open question that nobody has studied and that I would not want to assert either way.

Now the finding that matters most. In seven of those eight studies where clinical judgment won, the clinicians had more data than the formula did. And when clinical interview data were available to the clinicians, mechanical prediction's advantage increased rather than decreased, which means the interview was making the experts worse.

Read that carefully, because it is the most useful sentence in this chapter. The condition under which the human beat the formula was not that the human had better judgment. It was that the human had more information. Where they had the same information, the formula won or tied almost every time. And one particular kind of extra information, the face-to-face impression, made things worse rather than better.

* * *

The design decision where this cost me something was scheduling.

A queue of jobs needs assigning: which crew, which day, in what order. There is data. There is distance, duration, skill, availability, priority, and travel time, and an optimizer can produce an assignment better than a human on every one of those dimensions in about a second.

The first argument against automating it was the one I have been making all chapter, and it was wrong in the way the rule is usually wrong. The dispatcher said a formula could not capture what she knew. That is true and it is not an argument, because Grove's hundred and thirty-six studies are full of people who said the same thing and were measured and were mistaken.

So the question became specific: what do you know that is not in the data? And the answer was concrete rather than mystical. She knew that one customer's site manager would not let a particular crew back on the property. She knew that a job listed as one day was going to take two because she had spoken to the site that morning. She knew one technician was covering for something at home and needed short days that week.

None of that was in any system. All of it was decision-relevant. That is Scott's metis and it is also, precisely, Grove's seven-of-eight condition: the human was winning because she had more information, not because her judgment was superior to arithmetic.

Which reframes the design problem entirely and correctly. The right move is not to leave scheduling manual and it is not to automate it. It is to get the extra information into the system where it can be got, and to leave the decision with the person for the part that cannot. Site restrictions became data. Duration corrections became data. The technician's week did not become data and should not have, and the optimizer now proposes and the dispatcher disposes, with her overrides recorded, because an override rate is the only instrument I know that measures how wrong a rule is from inside the system running it.

The version of the rule I had been carrying would have stopped at the dispatcher's first answer and left the whole thing manual, and the schedule would have stayed worse than it needed to be for years on the strength of a claim nobody tested.

* * *

The boundary of the boundary is the comparison itself.

Where a data-based decision will not produce a genuinely better result, leave it manual. The entire content of that instruction is in an assessment of what a data-based decision would produce, and in practice almost nobody makes that assessment. The rule gets invoked before the comparison, as a reason not to run it, which inverts the sentence completely: it becomes where I believe a data-based decision would not be better, do not find out.

The second boundary is cost, and this is where the rule is strongest and I have undersold it. Grove's near-half of ties is a real result, and a tie means the formula loses, because the formula has to be specified, built, validated, maintained, and kept from going stale, while the person is already employed. The chapter on defining every variable priced that maintenance and the chapter on the same way every time showed how quietly a rule decays. A decision that gets made eleven times a year, by somebody who is there anyway, at a quality a formula would merely match, should stay manual forever, and that is most decisions in most organizations.

The third boundary is the one the weekend sits in, and it gets named separately rather than smuggled in. Some things are left manual because structuring them would cost something the person values, and the value is not accuracy. That is a legitimate reason and it is not the reason my sentence gives. My sentence claims the automated version would not be better. The honest version is that the automated version might well be better on every measurable dimension and I do not want it, because the wanting is itself one of the terms.

That distinction has a practical edge and it is not only about weekends. When somebody in an organization resists an automation, the first question is usually whether they are right about the accuracy, and that is the wrong first question. The better one is which of the two claims they are making. This will produce worse decisions is testable, often quickly, and being wrong about it is no disgrace, since Meehl's whole point is that experienced people are poorly calibrated here. This will cost me something I value is not testable, is frequently true, and is a different conversation with different people in the room.

Running those together is how automation arguments become unpleasant. The person making a preference claim gets answered with accuracy data, which does not address what they said, so they restate the preference in accuracy language, which is then refuted, and now they are both wrong and unheard. I have run that meeting from the wrong side of the table more than once.

* * *

Three things this rule cannot see.

It cannot see its own base rate. Stated as it is, the sentence sounds like it governs a large and important class of decisions. The best evidence available says that where the comparison has been run, the human wins substantially in about six percent of cases. The claim is true, and it is true about a small minority, and the sentence carries no hint of that proportion.

It cannot distinguish I prefer it this way from the method fails here. Both come out of the mouth as the same sentence, both are sometimes legitimate, and only the second is a claim about the world. This is the distinction the standardization chapter asked for and it belongs here: preference is a real reason and it is a different reason, and running them together is what makes the rule unfalsifiable.

And it has no account of the decision to stop looking. Every instance of the rule is a decision not to pursue something, which means it produces no artifact, generates no measurement, and leaves no trace that can later be audited. Every other rule in this creed can be checked against what it built. This one can only be checked against what it prevented, and nothing keeps that record.

The missing artifact is cheap to create and I have never seen anybody create it, including me until recently. A list of decisions deliberately left manual, one line each, with the date, the reason in the two-category form this chapter ends on, and the name of whoever made the call.

That list does three things nothing else does. It makes the invocation visible, which raises the cost of invoking it carelessly, because writing down left manual because the dispatcher knows things the system does not invites somebody to ask what things. It makes the decisions reviewable on a schedule, and the reason expires: information that could not be captured in 2023 is frequently capturable now, and nobody revisits a decision that left no record of having been made. And it converts a diffuse cultural tendency into a countable thing, so that an organization leaving everything manual can find that out rather than experiencing it as a series of individually reasonable choices.

The list is also the only defense I know against the asymmetry named earlier: every other ism in this creed generates evidence of itself, so the creed as a whole systematically over-reports its own building and under-reports its own restraint. A page of manual decisions is the counterweight, and it costs about a minute a month.

* * *

Here is the case against, and it is the strongest empirical case against anything in this book.

Meehl's argument, restated in its harshest form: expert confidence in unaided judgment is nearly uncorrelated with expert accuracy in unaided judgment, and people who make judgments for a living are systematically wrong about how good those judgments are. Seventy years of replication have not dented it. A hundred and thirty-six studies later, the formula wins substantially about half the time, ties about half the time, and loses in six percent of cases, most of which it loses because it was handed less information rather than because it reasoned worse.

Applied to the rule, that says something uncomfortable. Where a data-based decision will not produce a genuinely better result, leave it manual describes a real situation that occurs much more rarely than anyone invoking the rule believes, and the invocation is made by exactly the population the research shows to be poorly calibrated about it, which is experienced people assessing their own judgment.

The objection gets worse when you put it next to the interview finding. Clinicians given face-to-face contact did not improve. The formula's advantage grew. Which means the thing people most confidently cite as the irreplaceable human contribution, the sense you get from being in the room, measurably degraded performance in the setting where somebody checked. Everything the dispatcher told me about what she could see and the system could not was the same claim, and it happened to be true, and I know it was true only because I asked her to name it item by item and every item turned out to be a fact rather than a feeling.

The last form of the objection is about who the rule serves. A boundary ism is the one piece of a creed about systematization that gives its holder permission to stop. That makes it structurally attractive in a way none of the other eleven are, and it will be reached for disproportionately, by everyone, including its author, in exactly the cases where the work would be hardest and the gain largest. A rule whose invocation is pleasant and whose refutation requires building the thing you did not want to build is a rule that will be invoked too often forever.

* * *

The evidence does not destroy the rule. It converts it from an intuition into a test, and the test comes out of the same data that looked so damaging.

Seven of eight. The condition under which a person beats a formula is that the person has decision-relevant information the formula does not have. That is not a mood, it is not seniority, and it is not the feeling of expertise. It is an inventory question, and it has an answer.

So where a data-based decision will not produce a genuinely better result, leave it manual becomes a procedure rather than a permission, and it has three steps.

Name what you know that is not in the data. Item by item, out loud, in specifics. Not I know these customers, which is unfalsifiable, but this site manager will not accept that crew, which is a fact with a truth value. If nothing survives that exercise, the decision should be converted, and the belief that it should not was the thing Meehl's clinicians were wrong about.

Then ask which of those items could be captured. Most of them can. That is the move I missed for years: the correct response to the human knows things the formula does not is usually to put those things in the formula, not to abandon the formula. Site restrictions are data. Duration corrections are data. What remains after that pass is the genuine residue, and it is smaller than the first list every time.

Then leave the decision manual only for the residue, and say which of the two reasons applies. Either the remaining information cannot be captured, which is Scott's metis and a claim about the world. Or it could be captured and structuring it would cost something you value more than the accuracy, which is a preference and is legitimate and is not the same claim. The weekend is the second. The dispatcher's knowledge of who needed short days that week was arguably the first, and I am still not certain.

Where the tradition holds and the meta-analysis does not touch it: cost. Half of Grove's studies are ties, and a tie means do not build it. Almost nothing in an ordinary organization is worth formalizing on a ten percent accuracy gain, and the rule does most of its real work in that band rather than in the exotic cases where human judgment is irreplaceable. That is a smaller and more useful claim than the one I have been making, and it survives everything in this chapter.

And Scott's warning stands untouched by any of it, because it is about a different failure. The German foresters were not wrong about lumber yield. They were wrong about what the abstraction dropped, and no accuracy comparison would have caught it, because the thing that failed was not measured until it collapsed. That is an argument for humility about the scope of a formula rather than about its accuracy, and it is the reason the residue matters even when it is small.

Those two failures want different defenses and it is worth separating them, since this chapter has been treating them as one. Against an accuracy failure the defense is measurement: run the comparison, count the errors, see which side is better. Against a scope failure measurement is useless by construction, because the quantity that will fail is the one nobody thought to measure, and the formula will show clean numbers right up to the collapse.

The only defenses available against the second kind are slow and unsatisfying. Keep somebody looking at the actual thing rather than the abstraction, occasionally, with no particular question in mind, which is exactly the activity every other rule in this book is designed to eliminate. Watch for the second-order effects on a longer horizon than the project's review cycle. And treat a formula that has never surprised anybody as a warning rather than a success, because a model that has stopped producing surprises has either captured its subject completely, which is rare, or stopped being compared against it, which is common.

I do not have a better answer than that, and the answer amounts to keeping a person in contact with the material after the system has removed their reason to be. That is the same conclusion three other chapters reached, arriving here by a fourth road.

* * *

What survives:

Before leaving a decision manual, name what you know that is not in the data, specifically enough to be wrong about. Capture whatever can be captured. Leave manual only what remains, and say which reason applies: the information cannot be captured, or it could be and you value something more than the accuracy. Both are legitimate. They are not the same, and only the first is a claim about the world.

The original was a permission and it was unfalsifiable, because genuinely better was assessed by the person who did not want to do the work, and the best evidence available says that assessment is wrong far more often than it is right. The revision is a procedure with an inventory step in it, and the inventory is what makes it checkable by somebody other than its holder.

Two things go forward. This rule produces no artifact, so nothing records the decisions it prevented, and every other rule in the creed can be audited against what it built while this one cannot be audited at all. And I have not resolved whether the weekend is a boundary or a preference dressed as one, though I am now fairly sure it is the second and that this is fine, which is a different position from the one I held at the start of this chapter.