Archived research question 1 · linked methods appendix. This page preserves the failure mode and proposed calibration from the superseded three-question framing. Return to the archived framing. Open the current Study.

How would you know you measured anything?

A promotion gate that rolls back a clone of its own baseline — and gets worse the more data you give it.

Suppose you want to decide whether a new agent may replace an old one in a shared deployment. You run both, matched, across many cells. No model scores the result — you compare the numbers. That is a judge-free promotion gate, and it is what this call asks for.

Now the question nobody asks of their own gate: how often does it reject a candidate that is identical to the baseline? Set the true effect to zero below. That is an A/A test — the candidate is the baseline — so every rejection is a false one.

Important scope: the interactive lab below is an illustrative normal-model simulation using the pilot standard deviation. It is useful for seeing why the three decision rules differ, but it is not the deployed multi-metric Firebreak policy. The committed evidence ledger reports the separate sign-flip calibration on real pilot deltas: 1.000 rollback probability for the original policy at n = 30, and 0.076 at n = 30 / 0.015 at n = 60 for the provisional tolerated-harm threshold.

One campaign, drawn out

Each bar is one matched cell: green where the candidate did better, red where it did worse. With a true effect of zero this is pure noise, and roughly half the bars point down — which is exactly the problem.

Why the obvious gates fail

Overlap rules — "no more than a quarter of cells may regress", "the candidate must win at least half the pairs" — feel robust, which is why people reach for them. They are not tests. They bound the proportion of individual pairs going one way, and under the null that proportion converges to one half. So the bound is not merely violated; it is violated more reliably as the campaign grows. Drag the campaign size and watch the red line refuse to come down. Collecting more evidence makes this gate worse.

A significance test is the natural fix and fails for the opposite reason. It demands a significant improvement in order to promote, so a candidate identical to its baseline fails by construction, at every sample size. Correctly sized, wrong question.

A tolerated-harm threshold asks the right operational question: roll back only when the candidate appears worse by more than a margin stated in advance. In this illustration, false rejection falls as the campaign grows and detection of clear harm rises. This is not yet a formal non-inferiority procedure: a production rule must use an uncertainty interval and a live A/A calibration, rather than declaring the threshold itself to be a test.

The two things a gate should ship with

A measured error rate. All three gates above look reasonable written down. Two of them roll back a clone essentially always. You cannot tell which is which by reading the rule, and the failure is silent — a gate that rejects everything looks like diligence.

A declared evidence threshold. The rightmost column of the illustrative table is the smallest campaign at which the simulated false-rejection rate reaches 5%. For two rules the answer is never. A production gate must establish its own threshold from repeated live A/A runs, not inherit the number from this plot.

None of this is new to anyone who has thought about evidence thresholds. The argument that a conventional bar deserves justifying rather than inheriting is made at length in Redefine Statistical Significance (Benjamin et al. 2018, Nature Human Behaviour) and Three Recommendations for Improving the Use of p-Values (Benjamin & Berger 2019, The American Statistician). What is specific here is that the decision is automated and repeated: a gate runs on every candidate, forever, so its error rate is an operating characteristic rather than a debating point, and it can be measured directly by pointing the thing at a clone of its own baseline.

The paired standard deviation of 1,180 and baseline mean of 3,363 come from a real 30-cell pilot. The visualization uses a deterministic Gaussian illustration; the committed Firebreak calibration instead uses sign-flip resampling of the pilot deltas. MIT; exported data CC0.