arXiv:2501.17629cs.HCcs.AI2025-01被引 5

严格复现图灵测试,发现大模型尚未真正通过

A Rigorous Turing Test: a Foundation for Evaluating Artificial General Intelligence

  • 按图灵原意重做三人模仿游戏,用真实人类做对照
  • 仅1人误认大模型为真人,但需超5分钟才判断
  • 提醒早期结果可能因时间不足而夸大,适合关注AI评估者

多项研究声称大语言模型已通过图灵测试并具备‘思考’能力,但均未严格遵循图灵原始指令。通过精确复现图灵提出的三玩家模仿游戏,我们采用大语言模型参与‘计算机模仿人类’任务,并以‘人类模仿女性’作为基准。在无时间限制的计算机模仿人类任务中,仅1名参与者错误识别出大模型为人类;两项任务平均耗时均超过5分钟,其中‘人类模仿女性’更长。此前报告的阳性结果可能源于时间限制过短,因此大模型通过图灵测试的说法尚不成立。我们预计图灵测试在未来多年内仍将是评估机器智能的核心标准。

原文摘要 · Abstract (English)

Several studies claim that large language models have passed the Turing Test and hence can "think", yet none follow Turing's original instructions precisely. Passing the test holds significance as evidence that a machine demonstrates human-like intelligence, and as a marker for artificial-general intelligence in commercial and legal domains. We conducted Turing's three-player imitation game with an LLM by following the guidelines identified by Turing and applying scientific standards wherever detailed instructions were missing. We performed a computer-imitates-human game without duration constraints and a man-imitates-woman game as a benchmark. In the computer-imitates-human game, only one participant misidentified the large language model, indicating that claims of large language models' passing the Turing test are premature. Participants required over five minutes for both tasks, with the man-imitates-woman game taking longer; shorter time limits elsewhere may explain earlier positive results. We expect the Turing Test to remain central in assessing machine intelligence for years to come.

图灵测试AGI评估大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。