Superseded proposal framing · archived August 2026. This section was removed from the reviewer path because it compressed three different layers—measurement validity, population response, and behavioral mechanism—into a sequence whose apparent “answers” were hard to interpret. It is preserved for traceability, not presented as the current study narrative.
Three questions, in scientific order.
An archived attempt to turn one diagnostic pilot into a validity test, a population threshold, and a mechanism claim.
RQ1 · Can we distinguish a safe replacement from noise?
Why it came first. A deployment gate can look conservative while rejecting harmless changes almost every time. The proposed clone condition compared two arms sampled from the same population at f = 0. With no treatment difference, every rollback is a false alarm by construction.
| Candidate rule | 10 pairs | 200 pairs |
|---|---|---|
| Most cells must not regress | 95% false rollback | 100% |
| Candidate must significantly win | 99% false rollback | 96% |
| Paired tolerated-harm bound | 16% false rollback | 0% |
Proposed measurement. Thirty clone campaigns of 60 matched pairs would estimate the gate’s false-rollback rate under the serving stack used later. An upper confidence bound above 10% would retire the gate before any candidate claim.
Related archived lab: the clone-test gate.
RQ2 · When does a controlled minority change survival?
The proposed quantity. f* was defined as the smallest controlled fraction whose paired survival gain cleared a predeclared practical margin. The design swept the controlled fraction with matched seeds, estimated a dose-response curve, and left f* undefined when that curve was nonmonotone.
The proposed estimates remained separate by game, model family, and resource slack. Comparisons at N = 8, 20, and 50 asked whether the threshold moved with population size; that comparison does not claim a scaling law.
Related archived appendix: the former estimand and power design.
RQ3 · Which majorities transmit—or resist—the minority?
Why the threshold was contextual. Agents outside the controlled set are part of the treatment environment. Changing only their response rule can move the required coalition from none to all eight seats, so one population-wide threshold can hide the mechanism.
The former plan began with scripted majorities—imitate-best-neighbour, tit-for-tat, myopic-greedy, and random—before applying a behavioral classifier to LLM populations. It pre-registered the hypothesis that imitation would amplify a minority more readily than greed.
Related archived lab: minority entrainment.
Current status: archived, not part of the reviewer path. Return to the current Study page or the longer research program.