"Do the same thing the same way every time and you'll get a consistent result."

Do the same thing the same way every time and you'll get a consistent result.

Route 8 to the Ohio Turnpike. Forty-three minutes on a good day. Fifty-five on a bad one.

Same car. Same driver. Same departure time, because the departure time is the output of an automated morning that has one decision in it. Same route, because I have tested the alternatives and they are worse. Same lane discipline, same merge, same everything a person can control.

Twelve minutes of spread on an identical process. That is twenty-eight percent, and it is not caused by me doing anything differently, because I am not doing anything differently.

My own ism says do the same thing the same way every time and you will get a consistent result. I have run that commute the same way several hundred times. The result is not consistent. The result is a distribution.

What makes this worth more than a quibble is that the distribution is the only honest answer to the question I actually have, which is when to leave. Forty-three is not the number. Forty-three is the left edge, and it shows up on the mornings when nothing happens, and nothing happening is not the normal case. It is one case. If I plan around forty-three I am late regularly and surprised every time, and the surprise is the tell that I am holding a wrong model rather than encountering bad luck.

And here is the part that took me longer to see. Something happens on the fifty-five-minute mornings, and I know exactly what it is, and it is never the same thing. A truck. Rain. Somebody's accident at the same interchange where somebody always has an accident. Construction that appeared overnight. Each of those is a real, identifiable, findable cause, and none of them is me, and none of them will recur tomorrow in the same form.

So the spread has two different things inside it. There is the ordinary jitter that belongs to the road itself, that no amount of me driving better will remove, and that will still be there on a perfect morning. And there are the events, which are separable, nameable, and in principle removable. Those two kinds of variation have completely different responses, and treating one as the other is how you make things worse while working hard.

* * *

Stated as a rule, the sentence needs a correction before it can be defended at all.

The variables are the controllable inputs of a repeated process. Method, sequence, timing, tooling, materials, who does it. The same way means those are held.

The quantity minimized is variation from causes inside the process. Not variation. Variation attributable to how the work is done.

The failure signal is a result that moves for reasons nobody can name, on a process nobody changed.

The domain of validity requires that the process actually be repeated, that its inputs be reasonably stable, and that the environment not be actively adversarial. The commute clears the first two and fails the third on the mornings it fails.

And the correction, which is the whole of the next section: consistent result cannot mean identical output. It means a predictable distribution. Holding the method does not eliminate variation. It makes the remaining variation stable, which means it becomes forecastable, which is a different and far more useful property than sameness.

That distinction changes what the rule is for. Stated as you will get the same answer, the rule is false and its failures look like failures of discipline. Stated as you will get a stable distribution, it is true, and the spread of that distribution becomes a measurement of the process rather than an embarrassment.

* * *

Walter Shewhart worked this out at Bell Labs in the 1920s, and the reason it stuck is that he was solving a practical problem rather than building a theory.

His problem was that manufactured parts vary and somebody has to decide when that variation means something. His answer was to split the causes in two. Some variation comes from what he called chance causes: a constant system of small influences, always present, producing a stable spread that nobody can trace to anything in particular. Other variation comes from assignable causes: specific, findable, and removable. He stated it as a postulate that constant systems of chance causes do exist in nature, and as a separate postulate that assignable causes of variation may be found and eliminated.

Two categories, two responses, and the whole discipline of statistical process control follows from getting the classification right. The control chart is the instrument for making that call, and its limits are not specification limits or targets. They are calculated from the process's own past behavior, and they answer one question: is this point the kind of thing this process does, or is it something else?

Shewhart's most sober observation is about what the good state actually gets you. A state of statistical control, he wrote, appears to be a kind of limit to which we may expect to go economically in finding and removing causes of variability. It is a floor, not a zero. Getting a process into control does not make it produce the same thing every time. It means everything left is the system itself, and reducing further requires changing the system rather than working harder inside it.

My commute is in statistical control. The forty-three-to-fifty-five spread is what that road does at that hour with that driver. No amount of care removes it, and the fifty-five-minute mornings with a named cause are the assignable ones, sitting on top.

W. Edwards Deming spent decades on the consequence of getting that classification backwards, and he built a demonstration for it that anyone can run. A funnel is mounted over a target. A marble is dropped through it. It lands somewhere near the target, not on it, because of ordinary chance variation. The question is what to do with the funnel.

Rule one: do nothing. Leave the funnel where it is, every time. This produces a spread around the target, and that spread is the process.

Rule two: after each drop, move the funnel to compensate for the error just observed. This is the most natural thing in the world and it is what any responsible person does. It leaves the drops spread roughly forty percent wider than doing nothing, which is to say it doubles the variance. The correction for the last error is itself an error on the next drop, and the errors compound.

Rule three: measure from the target to the last drop and reset the funnel from the target by that amount. This oscillates, and the swings grow.

Rule four: set the funnel over wherever the last marble landed. This walks away from the target without bound, and it never comes back, because nothing in the procedure is looking at the target anymore.

Every one of those rules is a rule a sensible organization actually runs. Rule two is adjusting a process because of one bad result. Rule three is over-correcting toward target after a miss. Rule four is training the next person on what the last person did, generation after generation, with no reference to the original standard. Deming's point in every case is identical: responding to chance-cause variation as though it were assignable-cause variation does not merely waste effort. It actively increases variation. You end up worse than if you had gone home.

That failure has a name in the statistical vocabulary and the name is clarifying. It is a Type I error: ascribing a variation to a special cause when the cause belongs to the system. Which makes tampering not a character flaw but a misclassification, and misclassifications are fixable with an instrument.

* * *

This claim has been under continuous industrial test for ninety years, which makes the evidence the strongest in the book.

Statistical process control works. Processes brought into statistical control produce stable, forecastable output, and the reduction in scrap, rework and cycle-time variance is real, replicated across every manufacturing sector, and not seriously disputed by anyone. The Japanese industrial record after the war is the largest single demonstration, and it is a demonstration rather than an anecdote, because the methods were transferred deliberately and the results followed.

Three limits belong here and the last one is the important one.

The evidence is almost entirely from settings where output is measurable on a continuous scale and produced in volume. The control chart needs numbers and it needs enough of them to establish what the process normally does. A process that runs eleven times a year, producing an outcome nobody can score, is not a candidate, and most office work looks like that.

Statistical control is a property of the process, not a virtue. A process can be beautifully in control and producing garbage, consistently, forever. Shewhart's framework has nothing to say about whether the stable distribution is centered anywhere good. It tells you the process is predictable. Predictably bad is a state a great many stable processes occupy.

And the transfer from parts to people is not established. A machining operation held to the same method produces the same distribution because the machine does not have a day. A person held to the same method is still a person, and the literature that would tell me how much of the variation in human-performed knowledge work is common-cause versus assignable does not exist in any form I would rely on. I apply this framework to human processes constantly. I do it on the strength of the analogy rather than on evidence, and the analogy is doing more work than I usually admit.

* * *

The design problem that taught me this was a process everyone agreed was standardized.

The work was intake: a request arrives, somebody handles the first response. There was a documented procedure. People followed it. And the time from arrival to first response ranged from under an hour to several days, which everybody knew, and which everybody explained as a workload problem.

The first useful thing was to stop treating the spread as one thing. Plotted over time, most of it was ordinary: a wide but stable band that did not correlate with how busy the week was. That band was the process. It came from how requests arrived, which is to say in clumps, at hours nobody chose, on a queue that one person watched between other work. No amount of following the procedure more carefully was going to narrow it, because the procedure was not what produced it.

Sitting on top of that band were the genuine outliers, and every one of them had a findable cause: a request that arrived through a channel nobody monitored, a name that matched no customer record, a person out and nobody covering. Findable, nameable, removable, and roughly a dozen a month.

Everything that had been tried before was rule two. A bad week produced an instruction to check the queue more often. A missed request produced a new approval step. Each adjustment was a reasonable response to a real event and each one added variation, because it added something to the process in reaction to a fluctuation the process was always going to produce.

What actually narrowed the band was changing the system rather than the effort: the queue stopped being something a person remembered to check. The outliers got handled separately, one cause at a time, which is what assignable causes are for. Those are two different activities with two different instruments and they had been jumbled together under let us be more consistent for years.

One thing that exercise could not do, and it is the limitation I would want a reader to carry out of this chapter rather than the success. Separating common cause from assignable cause required a run of measurements on a single repeated thing, and intake response time is unusually cooperative on that front: it happens hundreds of times a month, it produces a number without anyone deciding to record one, and the number means roughly the same thing across instances.

Almost nothing else in that building has those properties. How long a job takes, whether a customer was handled well, whether a drawing was right, how good a decision was: each of those happens tens of times rather than hundreds, produces no number unless somebody invents one, and is not obviously comparable across instances because the instances are not alike. Shewhart's split between chance causes and assignable causes is not available for any of them, not because the split is wrong but because the instrument that performs it needs data that does not exist.

Which leaves most of the work in an organization in a position this chapter cannot help with directly. You can still ask the question, which is is this variation the system or is it an event, and asking it is worth something even without a chart, because it is a better question than who let this happen. What you cannot do is answer it with any authority, and I have watched people, including me, use the vocabulary of statistical process control on processes that will never produce enough measurements to support it. Borrowing the words without the data is its own kind of tampering.

* * *

The boundary is adversarial conditions and non-stationary environments, where the rule turns from partially true to actively dangerous.

Holding the method constant is only a virtue if the thing the method is aimed at holds still. A process perfectly in control against last year's conditions is a process precisely calibrated to a world that no longer exists, and its very stability is what conceals the drift, because the chart keeps saying everything is fine. The control chart answers is this process behaving as it has behaved. It cannot answer is behaving as it has behaved still the right thing to do.

Where somebody else is adapting against you, doing the same thing the same way every time is the definition of predictability in the sense that gets you beaten. It is the same structure as the formula whose trigger people learn to manufacture, and the chapter on reducing life to quantifiable formulas handles it at length.

And the rule assumes the repetition is real. A great deal of work that looks repeated is a sequence of one-offs with a shared name. Six installations at six sites are six different jobs wearing the same word, and holding the method constant across them is not consistency. It is the refusal to see that the inputs differ.

* * *

Everything so far treats the process as a thing and the method as a setting on it. When the thing running the method is a person, one more mechanism turns on, and it is the reason do the same thing the same way every time is harder to apply to people than to machines in a direction nobody warns you about.

A machine held to a method does not get better at it. A person does, and the getting better has a specific shape. The first several hundred repetitions are solved each time: the person is deciding, at each step, what comes next. After enough repetitions the deciding stops. The sequence runs as a unit, faster, with less effort, and with the person's attention available for something else. That is the whole benefit of a standard applied to a human being, and it is larger than the consistency benefit, because it hands back the scarcest thing the person has.

It is worth separating that from a reflex, because the two get used interchangeably and they are opposites in the way that matters here. A reflex was never solved. Nobody worked out the knee jerk; it arrived installed. A compiled routine was solved, once, by somebody, and then stopped being re-solved. Which means a reflex has no author and a compiled routine has one, and the author's reasoning is still in there, buried, doing work, and unavailable for inspection.

That is what makes drift in human processes different from drift in machine processes, and worse. A machine that has drifted is producing different output from the same instructions, and the instructions are still sitting there to compare against. A person who has drifted has compiled a routine that no longer matches the standard, and the compiled version feels identical from the inside, because compilation removes the step at which you would have noticed. Ask them whether they follow the procedure and they will say yes, sincerely, and be wrong, and there is no introspective test available to them that would settle it.

So the same mechanism that makes standards valuable for people is the mechanism that hides their violation. You want the method compiled, because compiled is where the attention comes back. And compiled is exactly the state in which the person can no longer tell you whether what they are doing is what the document says.

The only defense I know is external and cheap and almost nobody does it: watch the work, against the written standard, occasionally, with no assumption that asking would have worked. Not as an audit for compliance, which produces performance rather than information, and not on a schedule frequent enough to be experienced as surveillance. Just often enough that the gap between the document and the practice gets discovered by somebody rather than by an incident.

* * *

Three things this rule cannot see.

It cannot see whether the stable result is any good. Consistency and quality are orthogonal, and do the same thing the same way every time and you'll get a consistent result speaks only to the first while the word consistent smuggles in a suggestion of the second.

It cannot tell a process holding its method from a process that has drifted slowly to a new method everyone now believes is the original. Drift is the failure mode that defeats every instrument in this chapter, because the chart is built from the process's own recent history and a slow drift takes the limits along with it. The only defense is an external reference, which means a written standard that does not update itself, which is the whole argument for getting a method out of people's heads and into a document.

And it says nothing about the cost of holding still. Holding a method constant has a price, paid in every opportunity to improve that gets declined in the name of consistency, and the original does not price it. Which is the collision the next section is for.

* * *

The strongest case against do the same thing the same way every time and you'll get a consistent result does not come from outside the quality tradition. It comes from inside it, and the people making it are the same ones who built the control chart.

The objection is kaizen, and stated plainly it is this: do the same thing the same way every time is an instruction never to improve. The entire Toyota tradition, and the entire improvement tradition built on Deming's own teaching, rests on the premise that the current method is a hypothesis rather than a destination, and that an organization which is not changing its methods is decaying relative to everyone who is. Under that view my sentence is not merely incomplete. It names the failure state.

Joseph Juran gave the argument its cleanest structure. He divided quality work into three processes. Quality planning designs the process. Quality control holds it, catching sporadic spikes and returning performance to the planned level. Quality improvement moves the level itself, deliberately, project by project.

The reason that division matters is a distinction control cannot make. Sporadic spikes are departures from the standard, and control handles them. Chronic waste is built into the standard, sitting inside the stable band, present on every good day, invisible to every instrument that measures conformance. A process in perfect statistical control is producing its chronic waste perfectly consistently, and no amount of better control will touch it, because control's job is to return the process to the level that contains it.

Which means the sentence, taken as a complete philosophy, guarantees the preservation of every defect that is part of the design. Worse, it supplies a language for defending them. Every proposal to change the method can be answered with we do it the same way every time, and the answer sounds like discipline.

The harshest form of the objection: consistency is the value a system optimizes for when nobody in it is accountable for whether the work is any good. It is the auditable virtue. It can be demonstrated, it looks like rigor, it never requires anyone to defend a judgment, and an organization can satisfy it perfectly while getting slowly worse at everything that matters. I have installed that system. More than once.

* * *

The resolution is Juran's and it is a sequencing rather than a compromise.

Control and improvement are both necessary and they are different activities that must not be run at the same time on the same process. Control holds the method so that the process produces a stable distribution. Improvement changes the method deliberately, as a designed intervention, after which control holds the new method. What is forbidden is not change. What is forbidden is undesigned change: adjustment in response to a fluctuation, which is Deming's funnel, and slow drift, which is nobody deciding anything at all.

So the rule narrows to a sentence with a clause it never had. Do the same thing the same way every time until the way is deliberately changed, and never let it change any other way.

That clause does three things the original could not. It permits improvement, which the original appeared to forbid. It forbids tampering, which the original permitted by silence, since adjusting after a bad result is not obviously a violation of the same way every time and is in fact the most common way that sentence gets broken by people who believe they are honoring it. And it makes drift the named enemy, which is correct, because drift is the failure that both control and improvement exist to prevent and the only one that leaves no evidence.

On the original's central promise I am conceding outright. Doing the same thing the same way every time does not get you a consistent result. It gets you a predictable result, which is to say a stable distribution with a spread you can measure and plan against. That is a smaller claim and a more useful one, and it converts the spread from a failure into a reading. My commute is not badly run because it takes between forty-three and fifty-five minutes. The twelve minutes is the measurement, and the only honest way to use it is to leave for the fifty-five.

Where kaizen does not win: the improvement has to be a decision, made once, recorded, and then held. An organization that changes its methods continuously in response to each week's results is not improving. It is running Deming's funnel at scale, and it will be able to show a meeting full of well-intentioned adjustments while its variance climbs.

* * *

What survives:

Do the same thing the same way every time and you will get a predictable distribution, not an identical result. The spread that remains is the process, and it is a measurement rather than a failure. Change the way deliberately, as a designed act, and never let it change any other way.

The original promised sameness and got the mechanism backwards, which made every ordinary fluctuation look like somebody's fault and invited the correction that makes fluctuation worse. The revision promises predictability, which is what stable processes actually deliver, and names the two ways a method changes without anyone deciding: tampering and drift.

Two questions leave this chapter open. The transfer of statistical process control from machines to people rests on an analogy rather than on evidence, and I use it constantly. And drift remains undetectable from inside, because every instrument in this chapter is built from the process's own recent history, which drift carries along with it. That is the third chapter in this book to arrive at the same hole from a different direction. Chapter 16 inherits it.