arXiv:2606.21844cs.CLcs.CY2026-06

测试大模型辨别真人与人机对话的能力,发现现有方法有明显盲区。

Inverse Turing Bench: Evaluating Language Models as Judges of Human vs. AI Dialogue

  • 用成对对话样本对比真人对话与人机对话,判断哪组是纯真人
  • GPTZero准确率达89.41%,但语义分析易受角色提示干扰
  • 为评估模型‘心智理论’提供新视角,适合研究人机交互的学者

随着AI系统融入在线空间,区分其与人类对话变得日益重要。我们提出Inverse Turing Bench,一个评估大语言模型及其他模型在多轮文本对话中辨别真人与AI能力的基准。该基准包含成对对话记录:一组为两人间真实对话,另一组为人类与AI的交互。任务是正确识别哪一组为纯真人对话。我们初步评估了若干模型,结果表明GPTZero、Claude Opus-4.6和GPT-5.5准确率分别为89.41%、77.92%和75.94%。结果表明,基于统计的方法存在语义盲区,而基于语义的方法易受角色提示影响。本工作将逆图灵测试视为探测大模型心智理论的工具,推动人机区分成为AI系统的关键能力。实时基准可访问https://huggingface.co/spaces/roc-hci/Inverse-Turing-Bench-Leaderboard(匿名保护)。

原文摘要 · Abstract (English)

As AI systems integrate into online spaces, differentiating them from humans in conversations is increasingly important. We present Inverse Turing Bench, a benchmark that evaluates LLMs and other models on their ability to differentiate humans and AI in multi-turn text. The benchmark provides a collection of paired dialogue transcripts, wherein one dialogue is between two humans and the other is between a human and an AI. The task is to correctly identify which dialogue is human-only vs. human-AI. We evaluated a preliminary set of models against this benchmark, and found that GPTZero, Claude Opus-4.6, and GPT-5.5 achieve the highest accuracy: 89.41%, 77.92%, and 75.94% respectively. Our results suggest that statistical approaches to detection have semantic blind spots, but semantic approaches are susceptible to persona-prompting. Our work speaks to the Inverse Turing Test as a probe of LLM theory of mind, and motivates human-AI differentiation as a critical capability for AI systems. Our live benchmark can be found at https://huggingface.co/spaces/roc-hci/Inverse-Turing-Bench-Leaderboard (anonymity preserved).

人机区分大模型评估逆图灵测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。