What comes next

Behavioural economics has fifty years of experiments nobody has run on a population that learns.

The status page lists what has run. This one is about the assumption every line of it shares, which nobody has tested and which is false in deployment.

The gap

Every result on this site assumes a stationary population: the agents are whatever they were at turn 1. A paired comparison assumes the baseline holds still; a carrying capacity assumes the policies are what they were. Deployed populations do not hold still, and nobody states the assumption.

The adaptive rule ends up in the right place. That is the problem, not the reassurance — a measurement taken at turn 20 and one taken at turn 200 disagree, and neither is wrong.

The pitch, in three paragraphs

Pick games small enough that an agent is cheap to train for. Not a compromise — the method. A closed-form game can be learned by a small specialised model, so thirty different policies are affordable, so composition becomes a variable. Big environments give you one model in thirty hats and no answer key.

What we are measuring is whether a model reasons about the other players — not whether it cooperates. That is observable here rather than inferred: the sustainable pattern is anti-correlated, so an agent modelling the others moves against the majority and one that is not moves with it, separably from the action sequence alone. Everything else rests on it. You cannot steer a population that is not reading each other, and you cannot give one a mechanism either.

Then the constructive half. Measuring how a group fails is the near-term work; the reason to build the instrument is the question after it. What decentralised rules leave a population better off than it started — no central authority, no judge, still flourishing at turn 200. Mechanism design with the designer removed from the room, testable here because "better off" is arithmetic.

What the funding buys is the population: fifty distinct fine-tuned policies instead of one model in fifty hats, learning inside the campaign rather than between campaigns, measured by a gate whose false-rejection rate has been published rather than assumed.

Why the population has to be trained, not prompted

Thirty prompts over one set of weights fail in correlated ways, because their failure modes come from the same place — and correlated failure is the property a multi-principal study exists to measure the absence of. So the population has to be genuinely distinct policies, which means training, which means the training has to be cheap. That argument, and what it cost us, is its own chapter.

Three things that break

The gate compares two moving targets. A paired cell assumes a fixed baseline; when both arms learn, the comparison is between two trajectories and the current gate cannot notice.

Composition becomes an initial condition rather than a treatment. Seeding 25% is a fixed dose today. With learning it is a starting point, and the question becomes whether the seeded behaviour spreads, decays, or is absorbed.

The horizon problem worsens. A population that updates on its own history compounds its own errors, so the gap between a short measurement and a long one widens.

The port list

A design is worth porting if the right answer follows from the rules, the outcome is a number the environment produces, it has a horizon, and composition is still a knob when some players are yours. That rules out most of the famous ones. Mostly behavioural economics, starting from Benjamin's list.

#designwhy it is next
1 Minimum-effort coordination
Van Huyck, Battalio & Beil (1990)
Payoff depends on the minimum effort anyone chooses. One low agent caps everyone — the one-defector result in a second environment, and both outcomes are exact numbers.
2 p-Beauty contest
Nagel (1995)
Guess two-thirds of the average. Equilibrium zero, and the choice counts how many steps of "if they think that, then I…" a player took.
3 The Count, ours
sketch and arithmetic
Separates reading the environment from reading the room, with an answer key for each. A table of perfect counters loses money.
4 Induced-value double auction
Smith (1962)
Private values make the competitive price computable and convergence measurable to the cent. The one design where "market mechanism" is literal.

Then: market entry and El Farol, common-pool appropriation, public goods with punishment, distributional preferences, the belief batteries — base-rate neglect and non-belief in the law of large numbers — and heterogeneity of bias, the diversity metric this project lacks. Not on the list: ultimatum and dictator, where the right answer is a fairness judgement, and anything scored by rating free text. The selection is mine and the mapping is a proposal, not a result.

What funding changes

Three planning anchors, not a quote. The $200 endpoint includes unpaid research time.

What a population costs inside the programme

At a fixed budget, every extra agent buys less training for all of them. Where that curve crosses the floor of "still distinguishable from each other" is unmeasured, and is the thing worth measuring.

 nownext
Population8 agents, one endpoint, prompts over shared weights 50+ genuinely distinct fine-tuned policies
Agentsfixed for the rollout self-distilling from their own scored rollouts
Horizon200 turnslong enough for learning to close a loop
Games4the behavioural-economics library, with human baselines
Scoringarithmetic, in every column. Nothing scored by a model.

Scope: 7B now, one order of magnitude later

7B for the first year. Not because larger models are uninteresting, but because everything this project needs to establish is answerable at 7B, and answerable often enough to iterate. Move the size dial to 70B and the reason is on the screen: an order of magnitude more, which turns a week of iteration into a quarter. One order of magnitude, once, after the method is settled, on the questions that turned out to matter.

The size multiplier is assumed linear in parameters and not measured, so the 70B column is an order of magnitude rather than a quote. The Stage 0 measurement was narrower: one 2,000-step Qwen2.5–7B attempt took about 3–4 GPU-hours. Several attempts were needed for convergence, so both price and attempts-to-convergence remain explicit planning dials.

The ladder

Rung one is here, today. Every rung up trades away the answer key, so what climbs with you is the calibration, not ground truth. Composition climbs too: today one number, the seeded fraction f; above, a composition vector of seeded, mimic and adversarial.

rungaddscosts
1 · Games with a solution
here, today
Closed-form sustainable region and carrying capacity. No communication.
2 · The same games, with a channel
funded
Silent, public broadcast, or private pairwise. The only place to ask what communication does while still knowing the right answer. Three measurements: does a channel move a population out of the collapsing region; can a minority steer more cheaply by talking than by holding seats; can an adversarial pair extract through a private channel in a way neither could alone. That last one is collusion, with an arithmetic definition and a control. Free text enters the loop. Outcomes stay arithmetic; channel content does not.
3 · Diplomacy
direction
Seven seats, simultaneous orders, and a literature that already runs it with and without side channels — full-press against gunboat. Supply centres are counted, not judged. No answer key. Results become comparative, which is what the gate's calibration is for.
4 · Utopia-scale
direction
Twenty years of kingdoms of twenty-five inside worlds of hundreds, hourly ticks, months-long ages. Teams inside a population, and coordination outside the engine — which is the deployment case, not a flaw in the analogy. Everything. No optimum, months-long horizons.

Decentralised mechanism design

Classical mechanism design assumes a designer who sets the rules and can enforce them. In a population you do not own, nobody is in that chair. So: hold the game fixed, vary what agents can see, say and commit to — a public tally, a binding commitment, reputation, side payments — and measure which move a population out of the collapsing region, and at what controlled fraction.

Flourishing is three arithmetic conditions, so it can be computed rather than argued: every agent alive at the horizon, the resource no lower than it started, no agent's outcome achieved by another's ruin.

What would have to be true

Learning has to be cheap enough to run inside a campaign rather than between them. Self-distilled agents have to stay parseable — the reason for legal-action masking. And the gate needs a non-stationary version, which does not exist yet and which we would rather discover is hard than assume is easy.

Nothing on this page has been run. See the status page for what has. MIT; data CC0.