A live benchmark for social simulation: models forecast the next public-opinion release before it is published. Locked, hashed, scored in public. No one picks the test set, because the test set does not exist yet.
Today a good simulator cannot prove it is good. Trust needs a source no one controls, not even us: a public test against the real future.
An interdisciplinary team from 14 institutions: computer scientists, AI engineers, cognitive scientists, social scientists, political scientists, and economists. No single person decides the rules.












Left of the green line, every number is already published. Entries are hashed and locked 48 hours before each release; past the line the answer does not exist yet, and the release itself grades everyone in public.
Each round: forecast the tracker’s next release. Entries lock 48 hours before the number comes out; the release itself grades everyone; then the next round starts.
forecasts/<round>/<you>.json
{ "topline": { "mean": 39.0, "sd": 2.0 },
"cells": {
"Democrat": { "mean": 5, "sd": 1.5 },
"Republican": { "mean": 87, "sd": 2.0 },
...16 cells: party, age, race,
gender, education } }
Season 0 already grades the mainstream models under every harness on the left. The teams on the right are invited through a sealed track that keeps their product private.
BASE MODELS · 14 in season 0
HARNESS FAMILY 1 · context: what goes into the window
HARNESS FAMILY 2 · elicitation: the prompting strategy
Semilattice
Swift Centre
YuLan-OneSim
UniPat AIThe sealed track: the arena never sees the product, only its forecasts, filed before each lock and scored in public.
Every round is one question about a real tracker’s next release. Some ask about the whole country, some about a specific public.