the soak
Every new rule looks reasonable on the day someone proposes it. Switch it on by decision and you find out whether it was tuned right the expensive way: it stops good work until people learn to route around it, or it waves through the exact thing it was built to catch. A soak replaces the decision with a measurement. The control runs where it records what it would have done, against work whose answers are already known, and it earns the right to act only when the numbers say so.
the shape
The move is small and it changes everything. While it soaks, the control computes its verdict and does not apply it. It watches real work, or a labelled sample, or a load profile, and writes down what it would have done. Nothing it decides reaches anyone.
Then you read the record against rules written before the run. The ordering is the whole discipline. Fix the workload and the pass mark first, because a bar moved once the numbers are in is not a bar, it is a preference. Pass, and the control graduates and starts acting, with its threshold set from the measurements rather than from a round number someone liked. Fail, and it keeps soaking while the old behaviour stands.
why bother
Three ways a well-intentioned control goes wrong, and all three are invisible until it is already live.
The gate fires on things that were fine. Overriding becomes reflex, and a check nobody reads still costs everyone time. You now pay for the gate and get no signal back.
the common failureSet the other way, it passes exactly the defects it existed to stop, while the green tick tells everyone a check happened. Worse than no gate, because it manufactures confidence.
the quiet failureImpose a new step by authority and you get compliance theatre. Run it beside real work first and the argument for it becomes evidence instead of seniority.
the human failureThe right sensitivity is a property of the data, not of anyone's intuition about the data. A soak is the apparatus that reads it off.
anatomy
Naming the parts is what lets a load test and a review-gate trial be recognised as the same discipline, run by different people for different reasons.
The thing being de-risked. A gate, a workflow, a system about to meet real load.
What it is allowed to do while soaking, written down in advance. Not a promise: a property proven by a test.
What it runs against. Known-good and known-bad cases, real traffic, or a load profile. Fixed and recorded before the run.
What you compute over that workload. Separation, false alarms, detection, latency, what grows over time.
The arithmetic that turns measurements into graduate or hold, and that sets the number the control will actually use.
The explicit, evidence-bearing flip to acting. Or the decision to hold, keep the old behaviour, and keep soaking.
Give a trial those six and it is a method. Leave one out and it is a hope with a start date.
three uses
One shape, three jobs. What changes is what you are afraid of.
Run it over cases whose answers you already know, good and bad. Measure how cleanly it separates them, then pin the threshold inside the measured gap instead of guessing a round number.
fear: the wrong sensitivityDrive the system at production intensity for long enough that slow failures surface: memory, file handles, latency drift, contention. The output is a degradation curve, not a green tick.
fear: it degrades over timeRun the new process advisory on real work. Make it required only once it has shown it catches real problems at a false-alarm rate the team will genuinely tolerate.
fear: nobody will keep itthe ladder
A control does not go from idea to blocking. It climbs, and each rung is bought with measurements from the rung below.
each arrow is a graduation, and each one is bought with evidence
It runs dark first, recording where nobody sees it, so the first measurements cost nothing and embarrass no one. Then advisory, where it shows its verdict but cannot stop anything, which is also how the larger sample gets collected. Then a canary, acting for real but only on a small reversible slice, because how well a decision separates good from bad is a different question from what happens when it starts biting. Only then does it block broadly.
Agreement is not correctness. Several reviewers concurring proves consistency, and three that share a blind spot will agree confidently and be wrong together. Consistency earns a control the right to be considered. Only measured validity against known outcomes earns it the right to block.
the limits
A soak is strong evidence, not a proof, and the difference matters most when the numbers look flattering. The one that catches people out is sample size.
False alarms on twenty-four known-good cases. A perfect run.
The false-alarm rate that run actually rules out, at 95% confidence. Not 0%.
Clean cases needed before you can honestly claim a rate under 5%.
a perfect score on a small sample is a small claim, stated precisely