Cameron R. Jones · 2024 · controlled online behavioral experiment · n=?

Does GPT-4 pass the Turing test?

Cited 59 times in the scientific literature.

Level 2 - randomized trial

Level 2 by design analogy (controlled behavioral experiment), not clinical CEBM.

OpenAlex W4401043418 · doi:10.18653/v1/2024.naacl-long.290 · record verified 2026-08-31

What was done

GPT-4, GPT-3.5, and ELIZA were evaluated in a public online Turing test alongside human baselines. Interrogators interacted with counterparts and judged whether they were human or AI. The study analyzed pass rates across models and prompts, the decision criteria reported by interrogators, and the relationship between participant background (knowledge of large language models and games played) and detection accuracy.

What was found

The best-performing GPT-4 prompt achieved a pass rate of 49.7%, exceeding ELIZA (22%) and GPT-3.5 (20%), but remaining below the human baseline rate of 66%. Evaluator decisions were driven primarily by linguistic style (35%) and socioemotional traits (27%). Greater participant knowledge about large language models and a higher number of games played were positively correlated with AI detection accuracy.

Why it matters

This study provides an empirical benchmark for naturalistic AI deception, showing that frontier language models can convince human interrogators roughly half the time, while demonstrating that human detection capability improves with familiarity and practice.

Limits

The abstract does not report sample sizes for either participants or total games played. The public online setting introduces potential self-selection bias and uncontrolled participant environments, and evaluations were restricted to brief conversational interactions.

Cited by