arXiv:2503.23674cs.CLcs.HC2025-03被引 103

GPT-4.5在人机对话中被73%的人误认成真人,首次通过标准图灵测试。

Large Language Models Pass the Turing Test

  • 让大模型模仿人类角色,进行双人同时对话测试
  • GPT-4.5有73%概率被误认为真人,远超随机水平
  • 首次实证显示大模型可通过标准图灵测试,适合关注AI智能边界者

我们对四个系统(ELIZA、GPT-4o、LLaMa-3.1-405B 和 GPT-4.5)在独立人群中进行了两次随机化、受控且预先注册的三人群体图灵测试。参与者同时与一位真实人类和一个系统对话5分钟,随后判断谁是真人。当被提示采用类人身份时,GPT-4.5 被判断为人类的比例达73%,显著高于对照组选择真实人类的比例。LLaMa-3.1 在相同条件下被判断为人类的比例为56%,与人类无显著差异;而基线模型(ELIZA 和 GPT-4o)的胜率分别为23%和21%,均显著低于随机水平。这是首个实证表明任何人工智能系统通过标准三人群体图灵测试的研究,对理解大语言模型所展现的智能类型及其社会经济影响具有重要意义。

原文摘要 · Abstract (English)

We evaluated 4 systems (ELIZA, GPT-4o, LLaMa-3.1-405B, and GPT-4.5) in two randomised, controlled, and pre-registered Turing tests on independent populations. Participants had 5 minute conversations simultaneously with another human participant and one of these systems before judging which conversational partner they thought was human. When prompted to adopt a humanlike persona, GPT-4.5 was judged to be the human 73% of the time: significantly more often than interrogators selected the real human participant. LLaMa-3.1, with the same prompt, was judged to be the human 56% of the time -- not significantly more or less often than the humans they were being compared to -- while baseline models (ELIZA and GPT-4o) achieved win rates significantly below chance (23% and 21% respectively). The results constitute the first empirical evidence that any artificial system passes a standard three-party Turing test. The results have implications for debates about what kind of intelligence is exhibited by Large Language Models (LLMs), and the social and economic impacts these systems are likely to have.

图灵测试大模型人类识别AI智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。