"Define every variable to its deepest identifiable condition."

Define every variable to its deepest identifiable condition.

There is an equation I run on Route 8 southbound, and I have stated it in this book already. If A plus C is less than B, I get the double cheeseburger and fries. A is the perceived nutritional cost of eating that instead of eating at home. B is the perceived value of the energy and time I get back by not cooking. C is what the purchase takes out of the week's eating-out budget.

Three variables. Now do what my own rule tells me to do and define each of them to its deepest identifiable condition.

Start with A. Nutritional cost of what, measured how? Calories is the obvious first move, and it is wrong within one step, because I do not care about calories on Tuesday at six in the morning. What I actually care about is how I will feel at two in the afternoon, which is some function of sodium and fat and how much of it I eat and what I ate the day before and whether I slept. So A decomposes into at least four things. Take sodium, since it is measurable. Sodium relative to what baseline? My baseline over what window, a day or a week? And the thing I am comparing against, eating at home, is not one thing either. It is whatever is in the house, which varies, which means A is not a number but a distribution over what groceries got bought, and groceries every couple of weeks is itself a term I have never defined.

I am four levels down, the variable has fanned out into a dozen, and the double cheeseburger question has not moved a millimeter closer to an answer.

So try B. The value of the energy and time reserved by not cooking. Time is easy and worthless: I know how long it takes to make breakfast. Energy is where it lives, and energy is not a quantity I have any instrument for. I could proxy it with hours of sleep, which I know. That proxy is wrong in a way I can name, because the mornings when I most want the drive-thru are not reliably the mornings I slept least. They are the mornings something is coming that I do not want to do, which means B is partly about the day ahead rather than the night behind, and I have just discovered a fourth variable that was never in the equation.

C at least looks like arithmetic. It is dollars. It is not arithmetic either. The cost of spending nine dollars is not nine dollars, it is whatever the nine dollars does not get spent on later in the week, and I do not know what that is at six in the morning because the week has not happened yet.

Every one of those three terms carries the word perceived in my own statement of the equation, and I put that word there years before I could have told you why. It is load-bearing. It marks the level at which I stopped digging.

Define every variable to its deepest identifiable condition. I have never once done it, including in the example I use to explain the creed.

* * *

Stated as a rule:

The variables are the terms inside another definition. Define every variable to its deepest identifiable condition operates one level down from the rule before it. Chapter 2 says a term has to be defined before people can act on it together. This one says the definition itself contains terms, and those terms contain terms.

The quantity minimized is the number of levels at which two people can silently differ. A shared definition that bottoms out in another undefined word has moved the disagreement rather than removed it. Covered meant three things; say covered means inside the warranty period and you have bought real ground, unless warranty period starts at shipment for one person and at installation for another, in which case you have bought nothing and now have a document saying you agreed.

The failure signal is a specification everybody signed and nobody can apply to a hard case. The easy cases come out right. The disagreement reappears at exactly the requests that matter, which are the ones the definition was written for.

The domain of validity is wherever the decision actually changes with depth. That clause is the whole chapter, it is not in the original sentence, and putting it there is what the rest of these pages are for.

The phrase doing the damage is deepest identifiable condition. Read one way it means: keep going until you hit something everyone can observe. Read another way it means: keep going. The second reading has no stopping rule, and for years I stated the rule in a form that permits it.

* * *

The most rigorous attempt to make this rule operational came from database design, and its history is the whole argument in miniature.

Edgar Codd published the relational model in 1970, and one of its requirements was that the values sitting in a relation be atomic. Not lists, not records with parts, not a field called address holding a street and a city and a postal code in one string. A column holds one indivisible thing. Getting a design to that state is the first step of normalization, and it does exactly what define every variable to its deepest identifiable condition asks: it takes every variable and pushes it down to its deepest identifiable condition, where identifiable means the database can identify it as a unit.

The payoff is real and it is not aesthetic. Once address is split, you can ask which customers are in Ohio without string matching. Once a repeating group is pulled into its own relation, adding a fourth phone number does not require altering the table. Every question you can ask is a function of how far down the decomposition went, and questions nobody anticipated become answerable for free. That is as strong a case for defining deeply as exists anywhere in engineering.

W. Ross Ashby had already supplied the formal floor, from a different direction, in 1956. His Law of Requisite Variety says that only variety can destroy variety: the number of distinct states available to a regulator has to be at least as great as the number of distinct disturbances it must handle. A thermostat with two settings cannot manage a building with six thermal zones. Not because it is badly built. Because it does not have enough distinct states to answer with.

Put that next to a definition and it becomes a lower bound on depth. If the world produces twelve materially different situations and the definition recognizes four, eight situations are getting an answer built for something else. The definition is not slightly imprecise. It is structurally incapable of distinguishing cases it will encounter, and no amount of care in applying it will help, because the distinctions are not in it.

So there is a floor, it is calculable in principle, and define every variable to its deepest identifiable condition points in the correct direction relative to it. Most definitions in most organizations sit below the floor. That is the honest case for the rule, and it is the last comfortable paragraph in this chapter.

* * *

The evidence for going deeper is thinner than the argument for it, and the thinness runs in a particular direction.

Normalization has a strong theoretical case and a well-documented practical limit. Database designers denormalize on purpose, routinely, because fully decomposed data requires reassembly on every read and the reassembly costs more than the redundancy saves. That is not a failure of the theory. It is the theory meeting a cost the theory does not price. The literature on this is enormous and its conclusion is consistent: the right depth depends on the read pattern, meaning on what questions get asked, meaning on the decisions the data serves.

Outside database design I could not find much. There is no body of research establishing that organizations which define their terms more deeply perform better, and the reason is probably that the study is close to impossible to run. You would need two organizations alike in every respect but definitional depth, a measure of depth that is not circular, and an outcome measure that is not itself a definition. Nobody has that.

What does exist is the requirements literature from the previous chapter, and it does not settle this question either. It establishes that errors in requirements are expensive when caught late. It does not establish that more finely decomposed requirements produce fewer errors, and there is a reasonable argument in the other direction, since a longer and more finely specified document is a larger surface for a contradiction to hide in and a heavier thing to keep current.

I want to be accurate about what I am doing here. I hold to defining every variable to its deepest identifiable condition, I have built systems on it for years, and when I went looking for the studies that would justify holding it strongly, they were not there. Codd's atomic values are an argument about what becomes answerable. Ashby's requisite variety is an argument about a floor. Neither is an argument that deeper is better without limit, which is what my sentence says.

* * *

The place I have watched this play out most clearly is a partner matrix.

The problem is simple to state. A job needs installing somewhere. Which partners can do it? The original answer lived in a spreadsheet with a column for each partner and a list of states, and the definition of can do it was a state abbreviation.

That definition is above Ashby's floor for the easy half of the work and hopelessly below it for the rest. A partner who works in Ohio does not work in all of Ohio. Their crews are in one metro and travel is billable past some radius. A partner licensed for one kind of work is not licensed for another. A partner with capacity in March has none in April. A partner who has done small jobs has not done a two-hundred-opening job. Every one of those is a distinct situation the world actually produces, and the state abbreviation answers all of them the same way.

So push down. Can do it decomposes into geography, and geography decomposes into a home base plus a radius plus what happens past the radius. It decomposes into licensure, which decomposes by jurisdiction and by work type and carries an expiration. It decomposes into capacity, which is a function of time and therefore never static. It decomposes into demonstrated scale, which requires defining what counts as having done a job of a given size, which requires defining size.

And here is where the rule turns on you. Every one of those decompositions is correct. Every one makes the matrix more accurate. And every one adds a field that some human being has to keep current, forever, about somebody else's business, using information that partner has no obligation to give you and that changes without notice. A licensure expiration date is a fact about the world that is true on the day it is entered and decays silently thereafter. Depth you cannot maintain is worse than shallowness you can, because a stale precise answer is trusted and a vague one is questioned.

The version that shipped went deep on geography and licensure, because those change slowly and can be verified, and deliberately stopped short on capacity, which changes weekly and cannot be verified without asking. Capacity is a phone call. The matrix says who to call.

* * *

The boundary is the level at which the decision stops changing.

That sentence is not in the original ism and it is the only thing standing between the rule and an infinite regress. Below some depth, further decomposition produces distinctions that are real, describable, maintainable, and irrelevant, because every branch leads to the same action.

Sodium content of a double cheeseburger is a real quantity. It is identifiable to several decimal places. It does not change what I do at the Howe Road exit, because there is no value of it that would make me drive past when the rest of the equation says stop, and no value that would make me stop when the rest says go. Defining it more precisely buys a number and no decisions.

Two failure modes live on either side of that line, and I have produced both. Stopping above it means the definition cannot tell apart cases that call for different actions, which is Ashby's floor and the state-abbreviation matrix. Going below it means paying maintenance forever on distinctions nobody acts on, which is a different way to be wrong and the more seductive one, because it looks like rigor and it generates artifacts that look like progress.

The test I now use is a question rather than a principle: is there a value of this sub-variable that would change what somebody does? If no, the level above it was the deepest useful condition, which is not what the original says and is what it should have said.

* * *

Three things the rule cannot see.

It cannot see its own maintenance bill. Definition is a one-time act and a permanent liability, and the original prices only the act. Every level of depth is a thing that can now be wrong, and a definition that has gone stale is more dangerous than one that was never written, because people act on it without checking.

That deserves more than a complaint, because it is the practical question the rule raises and does not answer: how do you find out that a definition has died?

Three instruments exist and none of them is sufficient. The first is to attach an expiry to the fact itself, which works wherever the underlying fact carries a date. A license expires. A certification expires. A contract has an end. Where a date exists, the system can surface the decay without anyone remembering to look, and this is the only one of the three that runs unattended. It covers a small fraction of what goes stale.

The second is the rate at which people override the definition. A partner matrix that says three firms can do a job, on a job where the scheduler calls a fourth, has produced an override, and an override is the system reporting on itself. A rising override rate on one field says that field's definition no longer matches the world. This works, it requires that overrides be recordable rather than invisible, and it is silent until somebody overrides, which means it reports on definitions people still consult and says nothing about the ones they have quietly stopped trusting.

The third is the arriving case that will not fit. A request comes in that the categories cannot classify, somebody forces it into the nearest box, and the forcing is the signal. This is the earliest warning of the three and the most reliably discarded, because forcing the fit resolves the immediate problem and nobody logs a near miss.

What none of them catches is the definition that is quietly wrong and still comfortable. It classifies everything, nobody overrides it, no case refuses to fit, and it has been pointing at the wrong quantity for two years. It generates consistent answers, which is exactly what it was built to do, and consistency is not evidence of correctness. It is evidence of consistency.

The chapter on reducing life to quantifiable formulas reaches the same wall from the other side, asking how you tell a formula that is still solving something from one whose purpose expired, and arriving at the same answer, which is that nobody performs the experiment that would settle it. Two chapters hitting one wall from opposite directions is the strongest evidence I have that the hole is real and not an artifact of how I framed either one. Chapter 16 inherits it.

A second thing the rule cannot see: it collides head-on with another of my own isms, the one holding that you should not make the specific unnecessarily vague. That rule is about not withholding detail the reader needs. This one, pushed to its limit, produces detail nobody can carry. Both are about the gap between what is said and what is needed, and they point opposite ways past a certain depth, and I do not have a general rule for where the crossing point is. Chapter 4 takes the vagueness side. Chapter 16 puts them in the same room.

And it has no account of definitions that are deep and wrong. Decomposition feels like verification and is not. You can push a variable down six levels, achieve perfect atomicity at the bottom, and have decomposed the wrong thing entirely, in which case the depth has only made the error more expensive to correct and more convincing to look at.

* * *

The strongest case against is not that deep definition is costly. It is that deepest identifiable condition names nothing.

The people who worked hardest on this got there first, inside the formal system where the concept was supposed to be exact. Codd's relational model requires atomic values. Christopher Date and Hugh Darwen, working within that tradition rather than against it, argued that the notion of an atomic value is ambiguous, and Date put it flatly: the notion of atomicity has no absolute meaning.

Their demonstration is the uncomfortable part. A character string is a single value and decomposes into substrings. A fixed-point number is a single value and decomposes into an integer part and a fractional part. An identifier like an ISBN is a single value and decomposes into a registration group, a registrant, a publication element, and a check digit, every one of which is meaningful and queryable. All of these are routinely stored in one field and treated as indivisible, and each of them has a legitimate context in which the right move is to split it. Atomicity is therefore not a property of the value. It is a decision about the value, made by somebody, for a purpose, and it has no reading independent of that purpose.

Date's own remark about this is the one that should worry me most: for many years, he writes, he was as confused as anyone else. If the concept resisted a lifetime of formal work by people who defined their terms for a living, my sentence has not solved it by asserting it.

Follow that through. If deepest identifiable condition has no absolute meaning, then define every variable to its deepest identifiable condition is not a rule. It is an instruction to keep going with no terminating condition, and the reason it has never destroyed anything I have built is that I have always stopped, using judgment I did not write down, at a level I chose for reasons the sentence does not contain. The rule has been getting credit for work done by something else.

And this is where Borges shows up, though the story is not quite his alone. The one-paragraph piece everyone quotes, published in March 1946 as Del rigor en la ciencia, was written with Adolfo Bioy Casares and appeared under a joint pseudonym, attributed inside the fiction to an invented seventeenth-century traveler. It describes an empire whose cartographers refine their craft until nothing will serve but a map of the empire at the scale of the empire, point for point. The map is perfect. It is also useless, and the following generations leave it to rot in the desert.

The reason that image has outlived its two hundred words is that it is not a joke about excess. It is a proof. A representation that omits nothing is the thing itself, and the entire value of a representation is in what it leaves out. Which makes define every variable to its deepest identifiable condition, read literally, an instruction to build the map at one to one.

* * *

The rule survives, and the thing it was missing was never depth. It was a reason to stop.

Here is the reason. Depth is set by the decision, not by the variable. A variable is defined deeply enough when no further decomposition would change what anyone does with it, and not one level further.

That has three consequences worth stating separately, because the original sentence hides all of them.

The right depth is not a property of the term. The same word gets defined to different depths in different systems, correctly, and neither is sloppy. Address is one field on an envelope and six on a tax form, and the tax form is not more rigorous. It is answering different questions.

The right depth changes when the decisions change. A definition that was correctly shallow becomes wrong the moment somebody starts asking a question it cannot answer, which is not a failure of the original work. It is what should happen, and it means definitions need review on a schedule rather than perfection at the start.

And the stopping rule is testable, which the original was not. Ask whether any value of the next level down would change the action. That question has an answer, and the answer can be checked by somebody other than the person who wrote the definition.

Ashby's requisite variety still holds on the other side, and the two bounds now meet. It sets the floor: the definition needs at least as many distinctions as there are situations demanding different responses. My revised rule sets the ceiling at the same place, approached from above. Below the floor, the definition cannot tell apart cases it will meet. Above it, every added level is maintenance without decision. The band between them is narrow, and it is the whole working space.

What I will not claim is that this makes the boundary easy to find. Knowing that the stopping rule is where the action stops changing does not tell you where that is for any particular variable, and finding it usually requires having gotten it wrong at least once. The partner matrix found its depth by shipping a version that was above the floor on geography and below it on capacity, and learning from which questions came back.

Which suggests how a first definition should be built, and it is not the way I used to build them. Go deep where the facts are stable and verifiable, because depth there is cheap to hold and the decomposition survives. Stay deliberately shallow where the facts move faster than anyone will update them, and make the shallowness explicit rather than accidental, so that the field says ask instead of quietly reporting a number from last quarter. A definition that admits what it does not know is more useful than one that guesses, because the first sends somebody to find out and the second stops the search.

* * *

What survives:

Define every variable to the depth at which further decomposition would no longer change what anyone does with it. That depth is set by the decision, not by the variable, and it moves when the decisions move.

The original said deepest identifiable condition and named no floor and no ceiling. The revision has both. The floor is Ashby's requisite variety, and it is a real constraint rather than a preference: a definition with fewer distinctions than the world produces will answer some situations with an action built for a different one. The ceiling is the decision test, and it is what stops the map growing to the size of the empire.

Two things go forward unresolved. The collision with don't make the specific unnecessarily vague is real and I have only named it. And the decision test assumes you know what decisions the definition will serve, which is exactly what you do not know when the definition is new, and the honest position is that first versions are guesses that get corrected by the questions they fail to answer.