Flockbench
Infrastructure for serving and studying mixed populations of language agents in behavioral-economics games. It records full traces and scores population outcomes with game arithmetic rather than an LLM judge.
Independent research · training · long rollouts · populations
A research program for cheap behavioral fine-tuning, experiments over long rollouts, and the study of mixed agent populations that must coordinate around shared resources.
Projects
Individual training, long-horizon evaluation, and population experiments are separate technical problems. These projects connect them without pretending that one prototype already solves the whole stack.
Infrastructure for serving and studying mixed populations of language agents in behavioral-economics games. It records full traces and scores population outcomes with game arithmetic rather than an LLM judge.
Infrastructure for generating, perturbing, and evaluating long agent rollouts. The research question is how training signals can reflect coordination, recovery, and drift that only become visible over time.
A theory-led study of partial control: when one actor controls some agents but nobody controls the whole population, what would it take to preserve both the community and its common resource?
Exploratory work on low-cost differentiable and progressive-denoising fine-tuning. It remains a research prototype and is not part of the current public release.