MIT LIVE

Social Simulation Arena

Social Simulation vs. The Real Future

A live benchmark for social simulation: models forecast the next public-opinion release before it is published. Locked, hashed, scored in public. No one picks the test set, because the test set does not exist yet.

social-simulation-arena.comSeason 0 · live since August 11, 2026
MIT
Why

Social simulation is easier to build than to trust

The work is real

  • Simile, Aaru and others sell simulated publics to enterprises and campaigns.
  • Gallup is testing simulated respondents against its own national panel.
  • New agents, personas, and synthetic publics ship every month.
  • The proof is missing

  • Every accuracy claim is graded by the team that makes it, in industry and in research alike.
  • So skeptics reject the whole field at once: Silver Bulletin bans AI polls outright. AAPOR and ESOMAR are writing rules.
  • Published benchmarks grade the past, where the answers already sit in training data.
  • Today a good simulator cannot prove it is good. Trust needs a source no one controls, not even us: a public test against the real future.

    MIT
    The consortium

    A neutral social simulation consortium

    An interdisciplinary team from 14 institutions: computer scientists, AI engineers, cognitive scientists, social scientists, political scientists, and economists. No single person decides the rules.

    MIT
    Stanford
    Harvard
    CMU
    UC Berkeley
    UChicago
    UWashington
    Johns Hopkins
    Northeastern
    UC San Diego
    Brown
    UBC
    NUS
    Microsoft
    MIT
    The idea

    A live benchmark, graded by the real future

    Left of the green line, every number is already published. Entries are hashed and locked 48 hours before each release; past the line the answer does not exist yet, and the release itself grades everyone in public.

    MIT
    How

    A live bench built on four recurring data sources

    Each round: forecast the tracker’s next release. Entries lock 48 hours before the number comes out; the release itself grades everyone; then the next round starts.

    SOURCES U. Michigan how America feels about the economy Economist/YouGov Trump approval, % + 16 demo cells Morning Consult Trump approval, registered voters % Generic ballot which party wins the House, in points ROUND the next release locks release − 48 h sha256 printed in public ONE SUBMISSION topline: mean 39.0 sd 2.0 cells: Democrat 5 · Republican 87 … 16 cells: party age race gender education RELEASE the tracker publishes BOARDS Overall Party Age Race Gender · Education Population shape Score: 0 = copying the last release · 100 = the oracle, a perfect forecast of the published number. Validated on a walk-forward replay of 323 real releases, 2015–2026. One file scores on every board; subgroup boards are views of the same locked submission, so no one can enter only the easy cells.

    Entering a round

  • Push. One JSON file per round, by pull request or the site’s one-click template, any time before lock (release − 48 h). CI checks the schema and the deadline and prints your sha256. Late means void, enforced in code.
  • Or API. Approve the question template and hand us a key; before each lock we call your endpoint and file the answer. Point-only APIs are fine: we sample a few times and build the distribution.
  • The format. A distribution, not a point: mean and sd per value, or quantiles. Confidence is scored, so bluffing costs. One file covers the topline and all 16 cells.
  • forecasts/<round>/<you>.json
    { "topline": { "mean": 39.0, "sd": 2.0 },
      "cells": {
        "Democrat":   { "mean": 5,  "sd": 1.5 },
        "Republican": { "mean": 87, "sd": 2.0 },
        ...16 cells: party, age, race,
           gender, education } }
    MIT
    Entrants

    Built for the teams that simulate society

    Season 0 already grades the mainstream models under every harness on the left. The teams on the right are invited through a sealed track that keeps their product private.

    What we test: base models × harnesses

    BASE MODELS · 14 in season 0

    OpenAI GPT-5.6 Luna · Sol · Terra
    Anthropic Claude Fable 5 · Opus 5 · Opus 4.8 · Sonnet 5
    Google Gemini 3.1 Pro · 3.6 Flash
    xAI Grok 4.5
    DeepSeek V4 Pro · V4 Flash
    Qwen Qwen3.8 Max · 3.7 Max
    Z.ai GLM-5.2

    HARNESS FAMILY 1 · context: what goes into the window

    1  zero-shot: training memory alone
    2  with-history: plus the last 10 releases of the series
    3  curated news digest: plus one news corpus, identical for every model
    4  live web search: open retrieval, same tool and budget for all

    HARNESS FAMILY 2 · elicitation: the prompting strategy

    1  direct ask: one prompt, one distribution
    2  persona sampling: ask N census-framed personas, pool the shares
    3  forecaster protocol: base rate, decompose, premortem, then commit

    Startups and teams already building this

    The sealed track: the arena never sees the product, only its forecasts, filed before each lock and scored in public.

    MIT
    Season 0 sample

    What a round asks: two sample questions

    Every round is one question about a real tracker’s next release. Some ask about the whole country, some about a specific public.

    Sample 1 · What does all of America think?

  • “Where does U.S. consumer sentiment land in the next monthly release?”
  • Series. University of Michigan Index of Consumer Sentiment: monthly, fixed wording since 1978, release dates scheduled in advance.
  • You submit. A national distribution: mean 55.0, sd 2.5, or quantiles.
  • Scored. The release prints, say, 53.4; CRPS scores the probability you placed near it.
  • Sample 2 · What does a specific public think?

  • “What share of adults under 30 approve of the President in this week’s release?”
  • Series. Economist/YouGov approval: weekly, topline plus 16 published demographic crosstabs.
  • You submit. That cell as a distribution, Under 30: mean 34, sd 3.0, with the other 15 cells in the same submission.
  • Scored. Cell by cell against the published crosstabs; subgroup boards read the same file.