Independent research · training · long rollouts · populations

Training agents for long-horizon coordination.

A research program for cheap behavioral fine-tuning, experiments over long rollouts, and the study of mixed agent populations that must coordinate around shared resources.

Formationlow-cost fine-tuning for behavioral-economics games
Long horizontraining signals drawn from behavior across extended rollouts
Populationsagents with different policies acting in one environment
Coordinationmechanical measures of common-resource survival and welfare

Projects

One research arc, four surfaces.

Individual training, long-horizon evaluation, and population experiments are separate technical problems. These projects connect them without pretending that one prototype already solves the whole stack.

Software

Flockbench

Infrastructure for serving and studying mixed populations of language agents in behavioral-economics games. It records full traces and scores population outcomes with game arithmetic rather than an LLM judge.

Early prototype

Freetimebench

Infrastructure for generating, perturbing, and evaluating long agent rollouts. The research question is how training signals can reflect coordination, recovery, and drift that only become visible over time.

Current study

Commons Game

A theory-led study of partial control: when one actor controls some agents but nobody controls the whole population, what would it take to preserve both the community and its common resource?

Coming soon

Continuous Judge

Exploratory work on low-cost differentiable and progressive-denoising fine-tuning. It remains a research prototype and is not part of the current public release.