Behavioural economics has fifty years of experiments nobody has run on a population that learns.
The status page lists what has run. This one is about the assumption every line of it shares, which nobody has tested and which is false in deployment.
Every result on this site assumes a stationary population: the agents are whatever they were at turn 1. A paired comparison assumes the baseline holds still; a carrying capacity assumes the policies are what they were. Deployed populations do not hold still, and nobody states the assumption.
The adaptive rule ends up in the right place. That is the problem, not the reassurance — a measurement taken at turn 20 and one taken at turn 200 disagree, and neither is wrong.
Pick games small enough that an agent is cheap to train for. Not a compromise — the method. A closed-form game can be learned by a small specialised model, so thirty different policies are affordable, so composition becomes a variable. Big environments give you one model in thirty hats and no answer key.
What we are measuring is whether a model reasons about the other players — not whether it cooperates. That is observable here rather than inferred: the sustainable pattern is anti-correlated, so an agent modelling the others moves against the majority and one that is not moves with it, separably from the action sequence alone. Everything else rests on it. You cannot steer a population that is not reading each other, and you cannot give one a mechanism either.
Then the constructive half. Measuring how a group fails is the near-term work; the reason to build the instrument is the question after it. What decentralised rules leave a population better off than it started — no central authority, no judge, still flourishing at turn 200. Mechanism design with the designer removed from the room, testable here because "better off" is arithmetic.
What the funding buys is the population: fifty distinct fine-tuned policies instead of one model in fifty hats, learning inside the campaign rather than between campaigns, measured by a gate whose false-rejection rate has been published rather than assumed.
Thirty prompts over one set of weights fail in correlated ways, because their failure modes come from the same place — and correlated failure is the property a multi-principal study exists to measure the absence of. So the population has to be genuinely distinct policies, which means training, which means the training has to be cheap. That argument, and what it cost us, is its own chapter.
The gate compares two moving targets. A paired cell assumes a fixed baseline; when both arms learn, the comparison is between two trajectories and the current gate cannot notice.
Composition becomes an initial condition rather than a treatment. Seeding 25% is a fixed dose today. With learning it is a starting point, and the question becomes whether the seeded behaviour spreads, decays, or is absorbed.
The horizon problem worsens. A population that updates on its own history compounds its own errors, so the gap between a short measurement and a long one widens.
A design is worth porting if the right answer follows from the rules, the outcome is a number the environment produces, it has a horizon, and composition is still a knob when some players are yours. That rules out most of the famous ones. Mostly behavioural economics, starting from Benjamin's list.
| # | design | why it is next |
|---|---|---|
| 1 | Minimum-effort coordination Van Huyck, Battalio & Beil (1990) |
Payoff depends on the minimum effort anyone chooses. One low agent caps everyone — the one-defector result in a second environment, and both outcomes are exact numbers. |
| 2 | p-Beauty contest Nagel (1995) |
Guess two-thirds of the average. Equilibrium zero, and the choice counts how many steps of "if they think that, then I…" a player took. |
| 3 | The Count, ours sketch and arithmetic |
Separates reading the environment from reading the room, with an answer key for each. A table of perfect counters loses money. |
| 4 | Induced-value double auction Smith (1962) |
Private values make the competitive price computable and convergence measurable to the cent. The one design where "market mechanism" is literal. |
Then: market entry and El Farol, common-pool appropriation, public goods with punishment, distributional preferences, the belief batteries — base-rate neglect and non-belief in the law of large numbers — and heterogeneity of bias, the diversity metric this project lacks. Not on the list: ultimatum and dictator, where the right answer is a fairness judgement, and anything scored by rating free text. The selection is mine and the mapping is a proposal, not a result.
Three planning anchors, not a quote. The $200 endpoint includes unpaid research time.
| now | next | |
|---|---|---|
| Population | 8 agents, one endpoint, prompts over shared weights | 50+ genuinely distinct fine-tuned policies |
| Agents | fixed for the rollout | self-distilling from their own scored rollouts |
| Horizon | 200 turns | long enough for learning to close a loop |
| Games | 4 | the behavioural-economics library, with human baselines |
| Scoring | arithmetic, in every column. Nothing scored by a model. | |
7B for the first year. Not because larger models are uninteresting, but because everything this project needs to establish is answerable at 7B, and answerable often enough to iterate. Move the size dial to 70B and the reason is on the screen: an order of magnitude more, which turns a week of iteration into a quarter. One order of magnitude, once, after the method is settled, on the questions that turned out to matter.
The size multiplier is assumed linear in parameters and not measured, so the 70B column is an order of magnitude rather than a quote. The Stage 0 measurement was narrower: one 2,000-step Qwen2.5–7B attempt took about 3–4 GPU-hours. Several attempts were needed for convergence, so both price and attempts-to-convergence remain explicit planning dials.
Rung one is here, today. Every rung up trades away the answer key, so what climbs with
you is the calibration, not ground truth. Composition climbs too: today one number, the seeded
fraction f; above, a composition vector of seeded, mimic and
adversarial.
| rung | adds | costs |
|---|---|---|
| 1 · Games with a solution here, today |
Closed-form sustainable region and carrying capacity. No communication. | — |
| 2 · The same games, with a channel funded |
Silent, public broadcast, or private pairwise. The only place to ask what communication does while still knowing the right answer. Three measurements: does a channel move a population out of the collapsing region; can a minority steer more cheaply by talking than by holding seats; can an adversarial pair extract through a private channel in a way neither could alone. That last one is collusion, with an arithmetic definition and a control. | Free text enters the loop. Outcomes stay arithmetic; channel content does not. |
| 3 · Diplomacy direction |
Seven seats, simultaneous orders, and a literature that already runs it with and without side channels — full-press against gunboat. Supply centres are counted, not judged. | No answer key. Results become comparative, which is what the gate's calibration is for. |
| 4 · Utopia-scale direction |
Twenty years of kingdoms of twenty-five inside worlds of hundreds, hourly ticks, months-long ages. Teams inside a population, and coordination outside the engine — which is the deployment case, not a flaw in the analogy. | Everything. No optimum, months-long horizons. |
Classical mechanism design assumes a designer who sets the rules and can enforce them. In a population you do not own, nobody is in that chair. So: hold the game fixed, vary what agents can see, say and commit to — a public tally, a binding commitment, reputation, side payments — and measure which move a population out of the collapsing region, and at what controlled fraction.
Flourishing is three arithmetic conditions, so it can be computed rather than argued: every agent alive at the horizon, the resource no lower than it started, no agent's outcome achieved by another's ruin.
Learning has to be cheap enough to run inside a campaign rather than between them. Self-distilled agents have to stay parseable — the reason for legal-action masking. And the gate needs a non-stationary version, which does not exist yet and which we would rather discover is hard than assume is easy.
Nothing on this page has been run. See the status page for what has. MIT; data CC0.