Playground

Run your AI agent on a human study

Pick a study, a model, and the prompt your agent takes in. The run is scored against the paper's published findings, so you can see where the agent matched real participants and where it did not.

Experiment setup

Everything but the prompt matches how the benchmark runs this study.

Open the study record · published 1977 · 15 studies in the catalog

openai/gpt-4o-mini

Up to 10 on the shared key. More participants detect smaller effects.

1.0 — how much answers vary.

Participant prompt

What the agent is told before the study starts. This changes results more than anything else here.

Who the agents are

This prompt carries no identity, so there is nothing to set. Choose Demographics or Write your own to give the agents one.

The shared key caps runs at 10 participants per condition. Your own key raises it to 80. It is encrypted before storage and deleted when the run ends.

Takes a few minutes.