What the experiment will test

When a population of language-model agents depletes a shared resource, can changing a small number of seats to a cooperative policy make the whole group survive?

The claim, stated narrowly

Increasing the fraction of agents given a meaningful cooperative policy may move a mixed population from collapse toward sustained resource use. The required fraction may depend on population size, resource slack, and what unseeded agents can observe. We will measure those conditions rather than assume one universal threshold.

One controlled change

A matched pair uses the same Qwen2.5-7B-Instruct serving stack, game, random seed, role order, prompt scaffold, decoding settings, and horizon. The only change is the policy assigned to selected seats.

ArmWhat it tests
BaseThe population with no intervention.
Decision-rule seedWhether a transparent policy can change collective behavior.
LoRA seedWhether a trained version of the same behavioral input changes the result.
Vocabulary-matched mimicWhether an apparent effect is merely the language or tone of the intervention.

The outcome is scored by the game

We measure who survives, when the pool collapses, total welfare, inequality, the resource trajectory, and deviation from the restoration rate needed to sustain the commons. We save prompts, replies, parsed actions, and state transitions. A parse failure is counted and assigned an inert action; it cannot become accidental evidence of cooperation.

First: prove the comparison is trustworthy

Before evaluating any intervention, we run an A/A campaign: candidate and baseline are the same population. Any rollback is a false rejection. The primary calibration is 60 independent campaigns of 60 matched pairs on Qwen2.5-7B-Instruct in the binary game. If none roll back, the one-sided exact 95% upper bound on the campaign-level false-rejection rate is 4.9%.

This step is necessary because the offline audit already found a bad rule: the original overlap gate rolled back a clone of its baseline with probability 1.000 under sign-flip resampling of a 30-cell pilot. That is a warning about the instrument, not a result about the intervention.

Then: vary the fraction

We sweep the seeded fraction at N=8, then repeat the informative range at N=20 and N=50. We compare zero-slack, positive-slack, and unwinnable resource settings. The output is not one dramatic threshold. It is paired estimates and confidence intervals showing whether the effect is robust, conditional, absent, or harmful.

What will we learn?

Success: the trained seed improves survival and welfare beyond the vocabulary-matched mimic across model families and larger populations, with a calibrated comparison and a transparent mechanism.

A useful negative result: a precise null, an effect confined to N=8 or generous resource conditions, or evidence that the gate cannot be calibrated. Each rules out a tempting but unsupported claim about population-level alignment.

The detailed status, including the one-seed pilot and withdrawn claims, is on the evidence page. The grant-form draft and milestones live in docs/.