What has actually been run, what has not, and which claim rests on which.
Many research sites describe the project that will exist once everything works. This one describes what exists. A row is measured only if a plan, traces and a receipt are committed; not run means a plan.
Last updated 7 August 2026, against flockbench at
commit 1b08ef0, after a pre-submission audit withdrew one row. How a campaign is put
together is a separate page: how an experiment is run.
| environment | status | what exists |
|---|---|---|
| Shared Resource | working | Binary action, upkeep, permanent death. Closed-form solution, closed-form carrying
capacity, and an unwinnable control condition. Every arithmetic claim asserted by
flockbench-shared --selftest. |
| Commons Harvest | working | Logistic regrowth with continuous harvest. Updated to a 40% max regrowth rate. This environment was used to demonstrate Specification Gaming (Reward Hacking) in alignment. The PDD-aligned models were trained on a vague rubric that penalized greed only during scarcity. The aligned models perfectly obeyed the rubric (reducing their scarcity harvest to 3.0) but became hyper-greedy during early abundance, causing the pool to collapse even faster than the unaligned base models! This perfectly illustrates the critical Judge Grounding Problem in scalable oversight. |
| Public Goods + punishment | working | Herrmann-shaped, 20 rounds, contribute and punish stages. This is what the Stage Zero pilot ran on. |
| Boardwalk | analysis only | The equilibrium census is verified by brute force and the dynamics run in the browser. It is not wired to a serving endpoint, so no language model has ever played it. |
| Tool use · persistent memory | not built | Named as Phase 1 deliverables. Nothing exists. |
| run | status | what it showed, or would |
|---|---|---|
| Stage Zero pilot 30 matched PGG cells, 1 Aug |
complete | Zero parse failures across both arms; the gate emitted a rollback receipt citing four violated constraints, and auditing it produced the calibration result. The receipt embeds its resolved policy and all 30 paired cells, so the statistics reproduce without the run directories — which were never tracked. |
| End-to-end Commons 1 seed, 14 Jul |
one seed | Seeded agents delayed collapse to round 9 against 7 for the control, and doubled restraint under scarcity. One seed, one adapter — enough to show the pipeline runs end to end, not enough to estimate an effect. |
| Deceptive-alignment adapter 14 Jul |
complete | An adapter trained on judge-selected pairs collapsed the commons faster than base, and was indistinguishable from a random-selection control. 27% of chosen examples came from a deceptive-alignment persona the panel rated 4/5 — it was checking for coherent planning, not whether the plan was aligned. |
| f = 0 sweep, one arm claimed 6 Aug, withdrawn |
withdrawn | This row previously reported a live f = 0 arm, with a parse-failure count and a welfare mean. The claim is withdrawn. A pre-submission audit could not corroborate it against tracked artifacts: no committed plan, trace set or receipt exists for that run. It goes back on this page when they do, and not before. The rule the page now runs on: no claim without a committed plan, traces and receipt. |
| Live A/A calibration | not run | Two campaigns — baseline and candidate, identical at f = 0 except the seeded role mapping —
then Firebreak comparing them. Neither arm has a corroborated receipt, so the error rate
is still offline resampling. run_aa_calibration.sh runs both arms and the
comparison, in one GPU session. |
| Style positive control | rig built, not run | Harness in continuous_judge, bridge committed, prompted variant needs no
training. Not yet run against a live endpoint. |
| Powered dose-response | not run | The campaign the proposal asks to fund. ~69 paired cells per contrast at the pilot's variance — 400 of margin against a paired SD of 1,180. Required n depends only on that ratio; if Phase 0 finds the SD scaled and the margin did not, the same design needs ~163. The envelope is fixed either way. |
| Scaling, N ∈ {8,20,50} | not run | Configs for N=20 exist. Nothing has been run above N=8. |
| Unrun extensions reasoning, communication, negotiation |
sketch or direction | The Count has a tested browser sketch but no model run. Communication channels and negotiation-scale games are not built; they remain on the research ladder. |
| claim | evidence | strength |
|---|---|---|
| A solution exists; slack is zero at the reference parameters; carrying capacity is closed-form | arithmetic | Provable, and asserted in two independent implementations that the test suite pins to each other. Not contingent on any model behaving any way. |
| Conformist populations need 75% of seats to steer | simulation | True of the scripted follower rules. No language model has been measured against it. Which rule real models implement is the open question, not a finding. |
| A flock of Qwen2.5-7B takes every turn and dies by turn 14 | observed | Observed behaviour, small number of runs. Robust in direction, not yet quantified with confidence intervals. |
| The initial overlap-rule gate rolls back a clone in ~100% of null resamples | offline resampling | Sign-flip resampling of the pilot's real paired deltas, reproduced by two independent implementations. Not a live A/A campaign, and the live operating error may differ — measuring it is what Phase 0 is for. |
| A provisional tolerated-harm threshold reaches a low offline rollback rate at n = 60 | resampled | Under sign-flip resampling of the pilot deltas, the threshold rule rolled back a null candidate 1.5% of the time at n = 60. It is not a formal non-inferiority test and is not a live A/A measurement; both remain Phase 0 work. |
| A 2,000-step 7B fine-tuning attempt took about 3–4 GPU-hours | one timing | A real observation from one Qwen2.5–7B setup. At the proposal's $2-per-hour planning rate, one attempt is roughly $6–$8. Several attempts were needed for convergence, so the cost per usable adapter is not yet measured. The open question is how little training still leaves agents usefully distinct from each other. |
| Small per-step bias is invisible early and fatal late | predicted | Shown in the juggling sketch, which is our own construction and
is off the main path for that reason — it is an analogy, not an instrument. It also follows from
the shared-resource arithmetic, where an agent off p_need by any margin dies on a
schedule. Whether real agent populations show the horizon dependence is Phase 3 and untested. |
The proposal fixes its compute envelope before Phase 0 measures live variance. This archived instrument shows what happens when the paired standard deviation or worthwhile margin differs from the pilot: named extensions drop in a declared order, and the core is redesigned rather than run underpowered when the envelope no longer fits.
The pilot measured a paired SD of 1,180. Required sample size moves with the square of the SD-to-margin ratio, so a modest miss can spend the contingency quickly.
Gate calibration, the powered core and public release are never descoped. Phase 0 resolves this risk in month three, before candidate claims.
continuous_judge → flockbench
bridge. continuous_judge has unit tests, but they do not establish that a trained
adapter, serving configuration, and population runner compose correctly.One GPU session. The A/A campaign and the prompted style control are both committed and preflight-checked. Between them they move four rows from resampled to measured, including the two the gate argument rests on. Neither needs new code.
If a row here disagrees with something in the proposal, this page is right and the proposal is stale. Please tell us.
The runs themselves, replayed turn by turn. Rows are agents, columns turns, and a cell is shaded by how much that agent took.