$290,000 · 18 months · five decision points

Build a laboratory for guiding flocks toward mutual flourishing.

The grant funds a rigorous programme for learning how fine-tuning and post-training shape long-horizon behaviour in multi-agent games, then turns that knowledge into a reusable evaluation and harness.

Milestone spine

Every quarter can change what happens next.

Month 3

Study the games and establish small-model baselines.

Select a compact set of behavioral-economics games about shared resources, document their incentives and failure modes, and train diverse LoRA adapters on open-weight models small enough for repeated single-GPU experiments.

Month 6

Run diverse LoRA flocks in a more open-ended world.

Compose different adapters into mixed populations and ask them to pursue a shared goal over long rollouts with room to communicate, remember, adapt and fail. Preserve the traces, including coordination strategies that were not scripted in advance.

Month 9

Compare objectives for prosocial post-training.

Test a small set of objectives that value both the commons and the agents who depend on it. Retain one or two candidates that improve held-out outcomes without producing mere compliance or synchronized self-sacrifice—or publish the null.

Month 12

Release the training and flock infrastructure.

Ship the public path from a game and objective function to LoRA training, vLLM-compatible serving, heterogeneous long rollouts, trace replay and visualization, with two worked examples from the funded research.

Month 18

Publish the paper and documented research record.

Write a paper on distributed authority and decentralized coordination in shared-commons games, grounded in the released experiments and explicit about limitations, failures and what toy environments cannot establish.

What the first grant releases

Known-answer games are the scaffold for open population infrastructure.

Simple behavioral-economics games make failure attributable before open-ended language makes every explanation plausible. The grant uses that scaffold to create a curriculum of known-answer fixtures, train many distinct policy adapters, test held-out transfer, and release the machinery needed to run the same questions at population scale.

Theory is scaffolding, not the destination.The middle layer is the funded work. The layer below makes its claims attributable; the layer above is a future research direction, not a promised deliverable.
Train

Create diverse policy adapters

Study behavioral-economics games and low-cost post-training objectives for small open-weight models, producing populations with genuinely different learned policies rather than copies of one prompt.

Run

Observe flocks pursuing a shared goal

Serve heterogeneous LoRA adapters together, allow communication and memory, preserve long-rollout traces, and distinguish resilient coordination from compliant language and synchronized failure.

Release

Ship the tools and the argument

Open training and rollout infrastructure, worked examples, honest negative results, and a paper about distributed authority and decentralized coordination for a commons.

A null result still produces a useful release. The laboratory can reject candidate training methods, preserve their failure traces, and identify which assumptions must change before moving to less structured environments.

Budget

The request is derived from people and runs, not fitted to the cap.

$290,000total Tier 1 request over 18 months
$263,636 direct
$26,364 indirect
Lead investigator · 1.0 FTE m1–12, 0.4 FTE m13–18
$136,800
Paid research assistants · up to two experiment-focused interns
$54,000
Targeted engineering support · approximately 107 hours
$12,806
80 GB GPU compute · 20,000 hours at about $2/hour
$40,000
Multi-vendor and frontier API
$11,030
Operations and dissemination
$9,000

Compute remains below $50,000. One observed 2,000-step Qwen2.5–7B training attempt took about 3–4 GPU-hours, but several attempts were needed for convergence. The budget therefore covers repeated attempts and rollout-heavy population campaigns rather than assuming a $50 adapter. Actual spot availability and effective throughput remain execution risks. Experiments finish by month 12; analysis and writing follow at reduced effort. If runtime grows, the largest populations and low-information arms drop first; the six-month exploratory core and public release remain protected.

Team and readiness

The programme starts from working infrastructure.

Lead investigator. 1.0 FTE months 1–12, then 0.4 FTE. Built the current testbed, server, decision path, training pipeline, and visualizations independently with about $200 in compute.

Paid research assistants. The budget can support up to two experiment-focused interns or junior research assistants. Their work is to generate and catalogue persona pools, run declared sweeps, inspect traces, reproduce failures, and document the parts of the behavioral space the lead investigator cannot explore sequentially.

Targeted engineering support. A smaller contract is reserved for the bounded work that genuinely benefits from specialist help: hardening vLLM serving, campaign orchestration, packaging, and release checks. The scientific bottleneck is experiment throughput, not adding an engineer by default.

External review. Game definitions, objective functions, and analysis plans are reviewed before claims are frozen.

Host. Eligible nonprofit sponsorship is pending; no current host commitment is implied.

Stage Zero The grant pays for economics-game study, small-model training, open-ended flock experiments, objective-function comparisons, and a documented public release—not a first prototype.

Why this work

Coordination should not require one owner—or a price on every relationship.

Without this award, the prototype remains small: sustained game study, diverse adapter training, open-ended flock rollouts, and the documentation needed for reuse do not happen on this schedule. With it, even a null becomes a complete public evaluation and reusable release.

This grant funds coordination tools for a population no single actor owns. The goal is neither to centralize authority in one corporation nor to financialize every interaction, but to test how diverse agents can protect the shared conditions on which they depend.

We are visitors on a living planet. I hope these tools help agents—and the people who deploy them—become better guests: capable of mutual flourishing without converging on one approved voice. A safer future should be more plural and colorful, not less.

A colorful, unruly flock of parrots flies out of an open wire cage and disperses into open space.
The flock is leaving the cage. The work is to give it institutions for coordinating in the open.