arXiv:2603.18981cs.LGcs.HC2026-03

让人类和大模型在共享空间中互为裁判与回应者,测试谁更像人。

Book your room in the Turing Hotel! A symmetric and distributed Turing Test with multiple AIs and humans

  • 构建分布式群组对话环境,人类与大模型同时担任评判者与被评者。
  • 17人19模型参与实验,部分大模型仍被误判为人类。
  • 平台支持移动端与电脑端统一接入,保障通信安全且可复现。

本文报告了基于大型语言模型(LLMs)与人类参与者混合社区的新型图灵测试实验——TuringHotel。经典的一对一图灵测试被重新定义为群体互动场景,其中人类与人工智能代理在限时讨论中相互扮演裁判与被测者角色。该社区通过UNaIVERSE平台(https://unaiverse.io)实现,形成一个定义角色与互动规则的“世界”,平台内置编程工具支持动态交互。所有通信通过认证的点对点网络进行,确保无第三方访问。平台提供统一界面,支持移动设备与笔记本访问,是实验成功的关键。实验包含17名人类参与者与19个大模型,结果显示当前模型仍可能被误认为人类。一些意外误判表明,人类特征仍有可识别性,但尚未完全明确。我们认为这是首次在分布式环境下开展此类实验,类似尝试对长期监测大模型演进具有国家层面意义。

原文摘要 · Abstract (English)

In this paper, we report our experience with ``TuringHotel'', a novel extension of the Turing Test based on interactions within mixed communities of Large Language Models (LLMs) and human participants. The classical one-to-one interaction of the Turing Test is reinterpreted in a group setting, where both human and artificial agents engage in time-bounded discussions and, interestingly, are both judges and respondents. This community is instantiated in the novel platform UNaIVERSE (https://unaiverse.io), creating a ``World'' which defines the roles and interaction dynamics, facilitated by the platform's built-in programming tools. All communication occurs over an authenticated peer-to-peer network, ensuring that no third parties can access the exchange. The platform also provides a unified interface for humans, accessible via both mobile devices and laptops, that was a key component of the experience in this paper. Results of our experimentation involving 17 human participants and 19 LLMs revealed that current models are still sometimes confused as humans. Interestingly, there are several unexpected mistakes, suggesting that human fingerprints are still identifiable but not fully unambiguous, despite the high-quality language skills of artificial participants. We argue that this is the first experiment conducted in such a distributed setting, and that similar initiatives could be of national interest to support ongoing experiments and competitions aimed at monitoring the evolution of large language models over time.

图灵测试多智能体大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。