Capraro · Scientific reports 2025 · computational benchmark / comparative AI evaluation study · n=?

A publicly available benchmark for assessing large language models' ability to predict how humans balance self-interest and the interest of others.

Level 5 - mechanism / opinion, no new human data

Computational benchmarking and evaluation of AI models using historical experimental datasets

PubMed 40595689 · doi:10.1038/s41598-025-01715-7 · record verified 2026-08-26

What was done

The authors developed a publicly available benchmark to assess the ability of large language models to predict human decisions involving trade-offs between monetary self-interest and the welfare of others. The benchmark compiles 106 textual instruction sets from dictator game experiments conducted across 12 countries alongside observed human choice data. The authors evaluated four commercial chatbots (GPT-4, GPT-4o, Bard, and Bing) on their capacity to predict human behavior across these tasks.

What was found

The abstract reports no exact numerical values or effect sizes. Qualitatively, none of the four evaluated chatbots met the benchmark standards. Bard and Bing failed to capture qualitative behavioral distributions. GPT-4 and GPT-4o successfully identified three major behavioral categories (self-interested, inequity-averse, and fully altruistic) but systematically underestimated the prevalence of self-interest and overestimated human altruism.

Why it matters

This study provides an open resource for auditing social cognition in AI and shows that current frontier language models hold an optimistic bias regarding human generosity, which could lead to flawed recommendations when models assist in social or economic decisions.

Limits

The abstract does not report specific quantitative error rates, confidence intervals, or the total number of human participants across the 12 countries. The scope is confined to stylized monetary dictator games and four specific model variants, limiting direct generalizability to other social dilemma frameworks, complex negotiations, or real-world non-monetary settings.

Cited by