Does GPT-4 pass the Turing test?
Level 2 - randomized trial
Level 2 by design analogy (controlled behavioral experiment), not clinical CEBM.
OpenAlex W4401043418 · doi:10.18653/v1/2024.naacl-long.290
What was done
GPT-4, GPT-3.5, and ELIZA were evaluated in a public online Turing test alongside human baselines. Interrogators interacted with counterparts and judged whether they were human or AI. The study analyzed pass rates across models and prompts, the decision criteria reported by interrogators, and the relationship between participant background (knowledge of large language models and games played) and detection accuracy.
What was found
The best-performing GPT-4 prompt achieved a pass rate of 49.7%, exceeding ELIZA (22%) and GPT-3.5 (20%), but remaining below the human baseline rate of 66%. Evaluator decisions were driven primarily by linguistic style (35%) and socioemotional traits (27%). Greater participant knowledge about large language models and a higher number of games played were positively correlated with AI detection accuracy.
Why it matters
This study provides an empirical benchmark for naturalistic AI deception, showing that frontier language models can convince human interrogators roughly half the time, while demonstrating that human detection capability improves with familiarity and practice.
Limits
The abstract does not report sample sizes for either participants or total games played. The public online setting introduces potential self-selection bias and uncontrolled participant environments, and evaluations were restricted to brief conversational interactions.
Cited by
- contradicts Artificial intelligence models such as GPT-3 passed the Turing test.