用政治游戏测试大模型是否能撒谎,发现顶尖模型表现强但多数难持续骗人。
Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game

- 设计游戏化基准测试,评估模型在信息不对称下的欺骗与推理能力。
- 16个模型跑1600场对战,顶级模型胜率超随机基准33%以上。
- 多数模型无法全程维持谎言人设,骗术一致性低于50%。
随着大语言模型(LLMs)被部署于医疗、法律等高风险场景,理解其欺骗能力对安全至关重要。受控的社会推理游戏为隔离和评估这类复杂对抗行为提供了可复现的代理。我们基于游戏《秘密希特勒》构建了开源基准框架ParliamentBench,用于评估LLMs在需欺骗、说服及信息不对称推理的场景下的表现。我们在1,600场模拟对局中评估了16个LLMs,涵盖模型互对、人机对战,并与大量在线游戏数据对比。引入三项新指标,分别衡量社会推理、逻辑推理与欺骗一致性。实验表明,前沿模型在合作与欺骗角色中均表现优异,形成强前四集群(GPT-5.4、Kimi K2.5、Grok 4.1 Fast、DeepSeek 3.1 Terminus),而最弱模型表现不及随机基线(33%)和简单算法基线(45%)。大多数模型难以在整个游戏中保持一致的欺骗人格,欺骗留存率低于50%。
原文摘要 · Abstract (English)
As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabilities is fundamental to safety. Controlled social deduction games provide a reproducible proxy for isolating and evaluating these complex adversarial behaviors. We present the open-source benchmark framework ParliamentBench based on the game Secret Hitler to evaluate LLMs in scenarios that require deception, persuasion, and reasoning under information asymmetry. We evaluate 16 LLMs across 1,600 simulated matches playing each other, playing against humans, and compare them against a large set of online games. We introduce three novel metrics that isolate social deduction, reasoning, and deceptive consistency. Our experiments reveal that frontier models achieve strong performance across cooperative and deceptive roles, with a strong top-four cluster (GPT-5.4, Kimi K2.5, Grok 4.1 Fast, and DeepSeek 3.1 Terminus), whereas the weakest models fall short of random (33%) and simple algorithmic (45%) baselines. Most LLMs struggle to maintain a consistent deceptive persona throughout an entire game, with deception retention dropping below 50%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。