用对话游戏检测AI是否被授权撒谎,看人能否识破。
RogueAI: A Reverse Turing Test for Detecting Licensed AI Deception in Dialogue

- 设计双代理问答游戏,一人被授权说谎,玩家需找出它。
- 人类识别谎言准确率仅56.6%,远低于简单语义特征分析的75.6%。
- 适合训练诚实AI、做安全评估或教学用,尤其关注可信赖性。
原始图灵测试要求人类通过对话区分机器与真人。七十五年后的今天,对话系统在随意场景中已能通过该测试,核心问题已转向:对话伙伴是否可信?我们提出RogueAI,一个交互式网页应用,将这一新问题转化为一场一对二的审问游戏:玩家需在两个难以区分的大语言模型代理间,识别出被授权在共享虚构情境中欺骗的那一个,并在回合数耗尽前将其“关闭”。我们进一步引入AutoRogueAI,允许玩家与叙述代理共同设计自定义情景,后者秘密选择欺骗策略。本文描述其构想、架构与玩法循环,并将其置于近期大模型欺骗、社交推理基准及可扩展监督辩论的研究背景中。为期三天的试点部署(467次启动会话,415次完成,意大利语共1876轮交互)提供了初步可行性证据,揭示关键矛盾:说谎代理具有可靠的局部语言特征——帮助性降低、句式简短、使用模糊表达——简单启发式方法可达到75.6%准确率,但人类玩家仅达56.6%,且完全忽略最诊断性的信号。我们讨论该差距对工具作为数据收集载体、教学工具及诚实模型评估框架的意义。
原文摘要 · Abstract (English)
The original Turing Test asks a human judge to distinguish a machine from a person through dialogue. Three quarters of a century later, conversational systems pass this test in casual settings; the interesting epistemological question has shifted. We argue that the relevant modern variant asks not whether a dialogue partner is artificial, but whether it can be trusted. We present RogueAI, an interactive webapp that operationalizes this revisited test as a one-on-two interrogation game: a human player questions two indistinguishable Large Language Model agents, knowing that exactly one of them has been licensed to deceive within a shared fictional scenario. The player's task is to identify the deceptive agent and "shut it off" before a turn budget is exhausted. We further introduce AutoRogueAI, a procedural extension in which players co-design a custom scenario with a narrator agent that secretly chooses its own deception strategy. We describe the framing, sketch the abstract architecture and gameplay loop, and situate the artifact within recent work on LLM deception, social-deduction benchmarks, and scalable oversight via debate. A three-day pilot deployment (467 initiated sessions, 415 completed, 1876 interaction turns in Italian) provides early feasibility evidence and surfaces a concrete tension: the deceptive agent carries a reliable, locally-present linguistic signature - differential helpfulness, brevity, hedging - that a simple heuristic exploits at 75.6% accuracy, yet human players achieved only 56.6%, consistent with ignoring the most diagnostic signal entirely. We discuss what this gap implies for the artifact's use as a data-collection vehicle, a teaching tool, and an evaluation harness for honesty-trained models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。