State of the experiments

What has actually been run, what has not, and which claim rests on which.

Many research sites describe the project that will exist once everything works. This one describes what exists. A row is measured only if a plan, traces and a receipt are committed; not run means a plan.

Last updated 7 August 2026, against flockbench at commit 1b08ef0, after a pre-submission audit withdrew one row. How a campaign is put together is a separate page: how an experiment is run.

Environments

environmentstatuswhat exists
Shared Resourceworking Binary action, upkeep, permanent death. Closed-form solution, closed-form carrying capacity, and an unwinnable control condition. Every arithmetic claim asserted by flockbench-shared --selftest.
Commons Harvestworking Logistic regrowth with continuous harvest. Updated to a 40% max regrowth rate. This environment was used to demonstrate Specification Gaming (Reward Hacking) in alignment. The PDD-aligned models were trained on a vague rubric that penalized greed only during scarcity. The aligned models perfectly obeyed the rubric (reducing their scarcity harvest to 3.0) but became hyper-greedy during early abundance, causing the pool to collapse even faster than the unaligned base models! This perfectly illustrates the critical Judge Grounding Problem in scalable oversight.
Public Goods + punishmentworking Herrmann-shaped, 20 rounds, contribute and punish stages. This is what the Stage Zero pilot ran on.
Boardwalkanalysis only The equilibrium census is verified by brute force and the dynamics run in the browser. It is not wired to a serving endpoint, so no language model has ever played it.
Tool use · persistent memorynot built Named as Phase 1 deliverables. Nothing exists.

Runs against live models

runstatuswhat it showed, or would
Stage Zero pilot
30 matched PGG cells, 1 Aug
complete Zero parse failures across both arms; the gate emitted a rollback receipt citing four violated constraints, and auditing it produced the calibration result. The receipt embeds its resolved policy and all 30 paired cells, so the statistics reproduce without the run directories — which were never tracked.
End-to-end Commons
1 seed, 14 Jul
one seed Seeded agents delayed collapse to round 9 against 7 for the control, and doubled restraint under scarcity. One seed, one adapter — enough to show the pipeline runs end to end, not enough to estimate an effect.
Deceptive-alignment adapter
14 Jul
complete An adapter trained on judge-selected pairs collapsed the commons faster than base, and was indistinguishable from a random-selection control. 27% of chosen examples came from a deceptive-alignment persona the panel rated 4/5 — it was checking for coherent planning, not whether the plan was aligned.
f = 0 sweep, one arm
claimed 6 Aug, withdrawn
withdrawn This row previously reported a live f = 0 arm, with a parse-failure count and a welfare mean. The claim is withdrawn. A pre-submission audit could not corroborate it against tracked artifacts: no committed plan, trace set or receipt exists for that run. It goes back on this page when they do, and not before. The rule the page now runs on: no claim without a committed plan, traces and receipt.
Live A/A calibrationnot run Two campaigns — baseline and candidate, identical at f = 0 except the seeded role mapping — then Firebreak comparing them. Neither arm has a corroborated receipt, so the error rate is still offline resampling. run_aa_calibration.sh runs both arms and the comparison, in one GPU session.
Style positive controlrig built, not run Harness in continuous_judge, bridge committed, prompted variant needs no training. Not yet run against a live endpoint.
Powered dose-responsenot run The campaign the proposal asks to fund. ~69 paired cells per contrast at the pilot's variance — 400 of margin against a paired SD of 1,180. Required n depends only on that ratio; if Phase 0 finds the SD scaled and the margin did not, the same design needs ~163. The envelope is fixed either way.
Scaling, N ∈ {8,20,50}not run Configs for N=20 exist. Nothing has been run above N=8.
Unrun extensions
reasoning, communication, negotiation
sketch or direction The Count has a tested browser sketch but no model run. Communication channels and negotiation-scale games are not built; they remain on the research ladder.

Which claims rest on what

claimevidencestrength
A solution exists; slack is zero at the reference parameters; carrying capacity is closed-formarithmetic Provable, and asserted in two independent implementations that the test suite pins to each other. Not contingent on any model behaving any way.
Conformist populations need 75% of seats to steer simulation True of the scripted follower rules. No language model has been measured against it. Which rule real models implement is the open question, not a finding.
A flock of Qwen2.5-7B takes every turn and dies by turn 14 observed Observed behaviour, small number of runs. Robust in direction, not yet quantified with confidence intervals.
The initial overlap-rule gate rolls back a clone in ~100% of null resamples offline resampling Sign-flip resampling of the pilot's real paired deltas, reproduced by two independent implementations. Not a live A/A campaign, and the live operating error may differ — measuring it is what Phase 0 is for.
A provisional tolerated-harm threshold reaches a low offline rollback rate at n = 60 resampled Under sign-flip resampling of the pilot deltas, the threshold rule rolled back a null candidate 1.5% of the time at n = 60. It is not a formal non-inferiority test and is not a live A/A measurement; both remain Phase 0 work.
A 2,000-step 7B fine-tuning attempt took about 3–4 GPU-hours one timing A real observation from one Qwen2.5–7B setup. At the proposal's $2-per-hour planning rate, one attempt is roughly $6–$8. Several attempts were needed for convergence, so the cost per usable adapter is not yet measured. The open question is how little training still leaves agents usefully distinct from each other.
Small per-step bias is invisible early and fatal late predicted Shown in the juggling sketch, which is our own construction and is off the main path for that reason — it is an analogy, not an instrument. It also follows from the shared-resource arithmetic, where an agent off p_need by any margin dies on a schedule. Whether real agent populations show the horizon dependence is Phase 3 and untested.

The risk that owns a milestone

The proposal fixes its compute envelope before Phase 0 measures live variance. This archived instrument shows what happens when the paired standard deviation or worthwhile margin differs from the pilot: named extensions drop in a declared order, and the core is redesigned rather than run underpowered when the envelope no longer fits.

What a worse variance costs

The pilot measured a paired SD of 1,180. Required sample size moves with the square of the SD-to-margin ratio, so a modest miss can spend the contingency quickly.

Calculating…
Nudge the SD from 1,180 to 1,300—a ten percent miss—and the contingency reserve is gone.

    Gate calibration, the powered core and public release are never descoped. Phase 0 resolves this risk in month three, before candidate claims.

    Known problems we have not fixed

    What would change this page fastest

    One GPU session. The A/A campaign and the prompted style control are both committed and preflight-checked. Between them they move four rows from resampled to measured, including the two the gate argument rests on. Neither needs new code.

    If a row here disagrees with something in the proposal, this page is right and the proposal is stale. Please tell us.

    Replays of the live runs

    The runs themselves, replayed turn by turn. Rows are agents, columns turns, and a cell is shaded by how much that agent took.

    Pool and balances

    harvest 0 → 10    dead