Three programmes named the same problem in the same year. This is an instrument for one corner of it.
Eight agents and a pool is a strange thing to build if the risks worth studying are large and networked. The argument for small is that this corner can be measured against a known answer, which is not true anywhere large.
Tomašev et al. (2025) take the patchwork hypothesis seriously: capability arriving through coordination among sub-AGI agents rather than in one monolith. Their answer is virtual agentic sandbox economies, with market mechanisms, auditability and oversight.
In their vocabulary: an impermeable sandbox economy, deliberately tiny. Nothing leaks into the human economy and nothing needs to, because the object of study is not the economy but the oversight mechanism. A sandbox you cannot calibrate is a simulation with a comforting name on it; this one has a closed-form sustainable region, so a rule can be pointed at a case whose answer is already known.
Their "differential access to tools, data, memory and resources" is the asymmetry the steering chapter is about: some seats are yours, most are not.
ARIA's thesis: coordination infrastructure — agents contracting programmatically, at scale, without intermediaries — is what preserves pluralism as agents get more capable.
The gate is a piece of that infrastructure in miniature. It decides, with no intermediary and no judge, whether a candidate agent may join a shared deployment — which is what "programmatic and without intermediaries" looks like in practice.
And the obvious version of the rule is wrong: in offline resampling it rolls back a clone of its own baseline in ~100% of null draws, and gets worse with more data. It looks reasonable written down. An automated rule with no published error rate is not trustworthy for being automated — it is confidently wrong at scale. Two things should ship with any such rule: a measured false-decision rate, and a minimum evidence threshold below which its outputs do not count.
The CAIF report names failure modes new in populations: collusion, conflict, destabilising dynamics, emergent agency, security. Three have a scale model here.
| failure mode | the model of it on this site |
|---|---|
| Destabilising dynamics | The commons collapses, and the collapse has an answer key. A population that dies at turn 14 in a game with a one-sentence solution is a measured failure, not an ambiguous one. |
| Conflict without resolution | The boardwalk at three vendors has no arrangement anybody will stop at. Steering there cannot mean arriving; it has to mean bounding a cycle. |
| Correlated failure | The least comfortable result here. The pattern this game needs is anti-correlated, so imitation — the social heuristic language-model agents most reliably exhibit — is close to the worst available rule. A population that copies each other locks into the state that starves it. |
Collusion and multi-agent security are not modelled here. Collusion needs a private channel between agents — at which point two adversarial agents extracting from a commons neither could break alone is a measurable event with an obvious control, the same pair with no channel. That is rung two of the ladder and is not built. Security is not on the ladder at all.
Read in order, one consistent thing: the interesting object was never a single agent.
| Huberman (1988) | The closest ancestor of this testbed. Computational agents sharing resources produce collective dynamics — oscillation, chaos — that no agent intends and none can see from inside. |
| Manheim (2018) | Multiparty dynamics and failure modes: Goodhart with more than one optimiser. |
| Drexler (2019) | Capability as services rather than as an agent — the earlier form of the patchwork hypothesis. |
| Critch & Krueger (2020) | ARCHES, and the multi/multi delegation case most alignment work skips. The case this project is in. |
| Dafoe et al. (2020) | Open Problems in Cooperative AI — the framing this call inherits. |
| Tomašev et al. (2025) | Virtual Agent Economies: the sandbox-economy framing, and the case for designing steerable agent markets on purpose rather than inheriting one. |
Also in the call's lineage and not summarised here: Minsky (1986), Wooldridge & Jennings (1995), Clifton (2020), Conitzer & Oesterheld (2023), Chan et al. (2025), Kolt (2025), and Hadfield & Koh (2025).
Four sections, and a stated preference for depth over breadth. This claims two, plus one bounded piece of a third, and says where it does not reach.
Five requirements are listed. Honest marks:
| requirement | where this stands |
|---|---|
| Scalable realistic agent numbers |
partial N ∈ {8, 20, 50}, and 50 is a real ceiling: the largest population for which the seed replication is also affordable. The threshold is not identifiable at N = 8, and the proposal says so. |
| High-fidelity frontier behaviour, not coarse abstractions |
partial Natural-language agents on a live endpoint with verbatim logs — but 7B, not frontier. The call welcomes "smaller, distilled models to serve as faithful proxies for frontier agents", which is the cheap-populations work in the funder's own words. Proxy fidelity then becomes a thing to measure. |
| Externally valid when to trust simulation-derived conclusions |
core What a closed-form solution buys: we can say what is claimed to generalise — the shape of a failure — and what is not, its magnitude. Transfer is a milestone rather than an assurance. |
| Safe and secure | yes Token economies in a sandbox. No tools with external effects, no network access from inside an environment, nothing that leaves. In the sandbox-economy vocabulary above, impermeable. |
| Reproducible | strongest Deterministic per configuration and seed, frozen trace schema, plan committed before the first scored run, and receipts that embed their own resolved policy. |
| A individual properties → system safety | central The whole steering question. How a population outcome depends on its members' dispositions — and the preliminary finding that it depends more on the update rule of the agents you do not control than on how many you do. |
| B vulnerabilities, adversarial sub-populations | by design, not yet run The composition vector carries an adversarial fraction, and resilience to it is a sweep the instrument supports. Attack propagation needs the channel that does not exist yet. |
| C emergence, phase transitions, scaling with population size | partial The collapse-to-sustain threshold is a transition in composition, and unusually it has a closed-form location to check against. Whether it moves with N is a milestone. |
| D theoretical foundations of collective agency | no Not attempted. We measure collective outcomes, not collective agency, and formal definitions of the latter are somebody else's contribution. |
| E dangerous emergent capabilities and goals | no Shutdown resistance, filter evasion, covert channels — none of it modelled, and claiming otherwise would stretch a resource game past what it carries. |
Not claimed as a section. Two of its four subsections have something specific in them.
| C multi-agent control and scalable oversight | small, and measured Firebreak is a control protocol for admission to a shared deployment. The unusual part is that its false-decision rate is measured and published, with a minimum evidence threshold below which outputs are void, and a deliberately broken candidate it must catch. |
| D mechanism and information design | direction, stated as direction The call's own words here are "(de)synchronisation … for stabilising volatile networks" and "agents designed to foster population-level cooperation and stability when reliance on centralised mechanisms is undesirable or infeasible." That is this project's result and its aim in a sentence each: the sustainable pattern is anti-correlated, so desynchronisation is the mechanism rather than a metaphor. Two channels are funded; the rest is direction. |
| A collusion and emergent communication · B attribution interfaces | not claimed Collusion becomes measurable once a private pairwise channel exists — rung two of the ladder, not built. Attribution across delegation chains is not attempted. |
Section 3, agent infrastructure, is not targeted. Saying so is part of answering a call that prioritises depth.
Not infrastructure, not a governance proposal. A measurement instrument for one question — how few agents you must control to hold a population you do not — in environments small enough that the answer is known before the experiment runs.
It aims at the constructive question those programmes are asking. Scaling Trust and the sandbox-economy papers both want mechanisms — rules under which a population coordinates well. Measuring has to come first, because a mechanism that sounds wise where nobody can say what winning was cannot be evaluated at all. The instrument is small so that the answer key survives.
The bet is that small and solved beats large and suggestive. In a rich environment you cannot tell a population that failed from one that faced an impossible problem. Here the carrying capacity is closed-form: a flock of eight survives exactly zero permanent defectors, and a measurement that disagrees is wrong.
What would make it not worth doing. If steering thresholds here carry no information about richer settings, the instrument is a curiosity. The mitigation is scale-up as a measurement rather than a hope: the same question at N = 8, 20, 50, across model families, with the disagreement reported either way.
Every characterisation above is mine and none of these authors has anything to do with this project. Where I have read only an abstract, I have said only what an abstract supports.