Playground

Run your AI agent on one of these studies

Choose a study and a base model, design the prompt your agent takes into the experiment, and see where it matched the original participants and where it did not.

Open the playground

HumanStudy-Bench

Agent Evaluations

See how AI agents reproduce human behavioral effects.

Read the HumanStudy-Bench paper →

Effect consistency

Illustrative model results

Loading effect size data…

Initial test suite

Leaderboard

Agent alignment across 12 reconstructed human-subject studies.

PASProbability Alignment Score
Measures whether agents and humans reach the same scientific conclusion at the inferential level.
ECSEffect Consistency Score
Measures how closely the magnitude and direction of agent effects match human ground truth.
Filter by Variant:
RankModelVariantPAS (Alignment)ECSCostTokensDetails
1gemini-3-flash-previewv3_human_plus_demo49.7%0.1593$2.78833,891,506Show
2gemini-3-flash-previewv4_background46.5%0.0076$5.27437,786,439Show
3gpt-5-nanov4_background45.9%0.0498$0.39195,197,263Show
4mistral-nemov3_human_plus_demo44.0%0.1311$0.44514,789,725Show
5qwen3-next-80b-a3b-instructv4_background43.4%0.1142$0.95274,695,595Show
6mistral-nemov4_background43.2%0.0389$0.20046,973,836Show
7mistral-nemov1_empty42.7%0.0700$0.39164,277,209Show
8gpt-oss-20bv1_empty41.9%0.0284$1.41588,450,730Show
9gpt-oss-20bv3_human_plus_demo41.8%0.0223$1.16777,561,478Show
10mistral-nemov2_human41.1%0.0339$0.38724,370,464Show
11ai-grok-4.1-fast-nonev3_human_plus_demo41.0%0.0030$0.84987,037,212Show
12gpt-5-nanov3_human_plus_demo40.1%0.0284$2.61359,988,192Show
13mistral-small-creativev3_human_plus_demo39.3%0.0124$0.45293,867,556Show
14claude-haiku-4.5v4_background38.9%0.0707$8.28195,970,093Show
15gpt-oss-20bv4_background38.8%0.0674$1.232011,628,944Show
16gpt-5-nanov2_human37.7%0.0078$2.879611,400,115Show
17deepseek-v3.2v4_background37.4%0.0249$3.04349,653,594Show
18gpt-oss-120bv3_human_plus_demo37.2%0.0557$1.74096,221,295Show
19gemini-3-flash-previewv2_human37.0%0.0962$2.76413,651,210Show
20gemini-3-flash-previewv1_empty36.8%0.1657$2.93553,678,290Show
21mistral-small-creativev4_background35.9%0.0832$0.63485,678,107Show
22gpt-5-nanov1_empty35.6%0.0650$6.504419,518,351Show
23qwen3-next-80b-a3b-instructv3_human_plus_demo35.1%0.1445$0.85903,843,866Show
24qwen3-next-80b-a3b-instructv1_empty34.9%0.1138$0.80903,386,421Show
25claude-haiku-4.5v3_human_plus_demo34.0%0.0213$6.46994,450,022Show
26gpt-oss-120bv4_background33.7%0.0074$0.99128,149,909Show
27deepseek-v3.2v2_human33.7%0.0124$0.80183,302,727Show
28ai-grok-4.1-fast-nonev4_background33.4%-0.0279$1.27367,232,883Show
29gpt-oss-120bv2_human33.3%0.0184$1.68626,130,775Show
30qwen3-next-80b-a3b-instructv2_human33.1%0.1798$0.82733,474,833Show
31gpt-oss-20bv2_human33.0%0.0036$1.36978,117,261Show
32ai-grok-4.1-fast-nonev1_empty31.9%0.0717$0.57845,291,271Show
33claude-haiku-4.5v1_empty30.4%0.0066$9.28774,737,817Show
34ai-grok-4.1-fast-nonev2_human29.9%0.0057$0.50125,140,245Show
35deepseek-v3.2v3_human_plus_demo29.7%0.0516$1.05223,753,460Show
36claude-haiku-4.5v2_human29.3%0.0586$10.16264,988,373Show
37deepseek-v3.2v1_empty29.3%0.0462$0.80003,253,924Show
38gpt-oss-120bv1_empty28.5%-0.0136$1.69396,049,728Show
39mistral-small-creativev1_empty25.9%0.0445$0.69754,905,549Show
40mistral-small-creativev2_human12.6%-0.0031$0.64224,482,800Show