Do LLMs pass the Turing test? And what does it mean if they do?

Do LLMs pass the Turing test? And what does it mean if they do?

🎙 Cameron JONES 👥 305 📅 November 13, 2025 ⏱ 92 min 👁 50 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

Turing testLLMGPT-4.5Llama 3.1human impersonation

Summary

Cameron Jones, an assistant professor at Stony Brook University, presents his research on whether large language models (LLMs) pass the Turing test. He begins by revisiting Turing’s 1950 paper, emphasizing the five-minute, three-party imitation game as the operational definition. He then reviews his series of experiments, from an initial open online test in 2023 to a controlled three-party study in 2025. In the latest study, GPT-4.5 with a persona prompt was judged human 73% of the time, significantly more often than actual humans, while Llama 3.1 with a persona achieved a 50% pass rate, at chance. The talk discusses the implications: passing the Turing test does not necessarily imply intelligence or language understanding, but it does indicate an ability to exploit superficial cues and impersonate humans. Jones also highlights the test’s value as a flexible, adversarial benchmark and its social consequences. He addresses objections, including the role of interrogator strategies and demographic factors, and concludes that while LLMs pass the test, this should not be overinterpreted.

166 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable empirical evidence on LLM performance in a Turing test, with a clear methodology and replication across studies. The argumentation is nuanced: Jones carefully distinguishes between passing the test and demonstrating intelligence, and he acknowledges limitations and alternative interpretations. He effectively uses examples and data to support his claims, and he engages with potential objections, making the argumentation solid.

Scientific Rigor, Source Quality, Title Accuracy

The presentation is scientifically rigorous, with references to the speaker’s own peer-reviewed and preprint papers. The methodology is described in detail, and the results are presented with appropriate statistical context. The title accurately reflects the content, and the talk does not overstate conclusions. The speaker also discusses the broader literature and debates, showing a balanced perspective.

133 words

Title / Content Match

The title accurately reflects the content: the talk presents empirical results on whether LLMs pass the Turing test and discusses the implications.

Quality & Reliability

8/10

The talk is based on peer-reviewed and preprint research by the speaker and collaborators, with clear methodology and transparent discussion of limitations. The speaker is an academic expert in psychology and AI. However, the presentation is an opinion/interpretation of results, and some claims are debated in the field.

Key Moments

Cited Sources

  • Large language models pass the turing test — Main paper presenting the three-party Turing test results
  • People cannot distinguish GPT-4 from a human in a Turing test — Conference paper on the two-party Turing test with GPT-4

Concurring Sources

Dissenting Sources

  • The Turing Test is not a good benchmark for AI — Some critics argue that the Turing test is not a valid measure of intelligence, which challenges the significance of the results.

Contribution & Novelties

The talk presents original empirical evidence that LLMs can pass a standard Turing test, with GPT-4.5 being judged human more often than actual humans. This challenges assumptions about AI’s ability to imitate humans and raises important questions about the nature of intelligence and human-AI interaction.

Pour aller plus loin :

  • Turing test - Wikipedia — Background on the Turing test and its interpretations.
  • Computing Machinery and Intelligence — Turing’s original 1950 paper.
  • GPT-4.5 - OpenAI — Information on the GPT-4.5 model used in the study.

85 words

Radar Profile

The radar profile shows high scores in information quantity, quality, and reliability, with a slightly lower technical level. This indicates a well-supported and informative presentation that is accessible to a broad audience.

Reliability 8/10

💬 No comments were provided for analysis.