arXiv:2503.06551cs.AIcs.CY2025-03被引 3

剖析ChatGPT-4图灵测试争议,提出两种有效测试范式与评估框架。

ChatGPT-4 and other LLMs in the Turing Test: A Critical Analysis

  • 区分绝对与相对通过标准,建立更精细的评估体系。
  • 证明三玩家与两玩家测试均有效,方法论意义不同。
  • 将测试建模为伯努利实验,强化统计分析基础。

本文批判性分析了Restrepo Echavarría(2025)关于ChatGPT-4未通过图灵测试的研究,指出其对测试实施严格性及数据有限性的批评缺乏充分依据。文章表明,三玩家与两玩家测试均为有效形式,各有独特方法论含义。研究提出区分绝对标准(机器误判概率≥人类正确识别概率)与相对标准(机器表现接近人类水平)的评估框架,实现更细致的评价。进一步地,通过将两类测试建模为伯努利实验(三玩家版本相关,两玩家版本独立),明确了理论通过条件的随机性本质,并强调实验数据需结合稳健统计方法解读。该工作不仅反驳了原研究的关键结论,还为未来客观衡量AI行为与人类相似度提供了坚实方法论基础。

原文摘要 · Abstract (English)

This paper critically examines the recent publication "ChatGPT-4 in the Turing Test" by Restrepo Echavarría (2025), challenging its central claims regarding the absence of minimally serious test implementations and the conclusion that ChatGPT-4 fails the Turing Test. The analysis reveals that the criticisms based on rigid criteria and limited experimental data are not fully justified. More importantly, the paper makes several constructive contributions that enrich our understanding of Turing Test implementations. It demonstrates that two distinct formats, the three-player and two-player tests, are both valid, each with unique methodological implications. The work distinguishes between absolute criteria for passing the test--the machine's probability of incorrect identification equals or exceeds the human's probability of correct identification--and relative criteria--which measure how closely a machine's performance approximates that of a human--, offering a more nuanced evaluation framework. Furthermore, the paper clarifies the probabilistic underpinnings of both test types by modeling them as Bernoulli experiments--correlated in the three-player version and uncorrelated in the two-player version. This formalization allows for a rigorous separation between the theoretical criteria for passing the test, defined in probabilistic terms, and the experimental data that require robust statistical methods for proper interpretation. In doing so, the paper not only refutes key aspects of the criticized study but also lays a solid foundation for future research on objective measures of how closely an AI's behavior aligns with, or deviates from, that of a human being.

图灵测试评估框架统计建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。