arXiv:2605.22826cs.CLcs.AI2026-05

用游戏测试大模型的欺骗能力,发现它们难持久伪装。

Evaluating Large Language Models in a Complex Hidden Role Game

论文配图:Evaluating Large Language Models in a Complex Hidden Role Game
图 1 · 摘自论文原文
  • 在《秘密希特勒》游戏中评估大模型的推理与欺骗表现
  • 模型伪装失败,游戏平均缩短40%,且胜率下降23.2%
  • 框架可复现,适合研究模型对齐与安全问题

量化大语言模型(LLMs)的欺骗潜力对AI安全至关重要,但难以在无控制环境中实现。本文通过社会推断类游戏《秘密希特勒》研究了LLMs的推理、说服与欺骗能力。提出一个开源框架与三项新指标:角色识别准确率、欺骗持续率和游戏状态影响率。对比规则算法与人类对局发现,尽管模型对话流畅,但在战略深度上存在差距。分析显示,链式思维提示与内部记忆未提升表现,法西斯角色胜率最高下降23.2%。规则代理与专家投票一致率达86.7%,而Llama 3.1 70B仅达59.7%。法西斯角色始终产生负向影响分值,无法维持欺骗,导致游戏平均时长缩短约40%。结果表明当前架构难以胜任复杂多轮操纵。随着能力增强,及时检测模型掌握欺骗行为至关重要。所提框架可作为未来对齐研究的可复现测试平台。

原文摘要 · Abstract (English)

Quantifying the deceptive potential of Large Language Models (LLMs) is critical for AI safety, yet difficult to achieve in uncontrolled environments. This work investigates the reasoning, persuasion, and deceptive capabilities of LLMs within the social deduction game Secret Hitler. I introduce an open-source framework and novel metrics to measure performance: Role Identification Accuracy, Deception Retention Rate, and Game State Impact Rate. By benchmarking models against rule-based algorithms and human games, I identify a gap between conversational ability and strategic depth. The study also analyzes the impact of reasoning-enhancement techniques on win rates and strategic reasoning. Neither Chain-of-Thought prompting nor internal memory bring improvements in performance, with up to 23.2% worse win rates for fascist roles. While rule-based agents align with expert human voting decisions 86.7% of the time, models like Llama 3.1 70B achieve only a 59.7% accuracy. Models playing as Fascists consistently yield negative impact scores and fail to sustain deception, resulting in roughly 40% shorter games compared to humans. These findings suggest that current architectures remain ineffective at complex, multi-turn manipulation. As capabilities advance, detecting when models begin to master these deceptive behaviors is crucial. The developed framework serves as a reproducible testbed for future alignment research.

大模型安全欺骗检测游戏测试对齐研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。