用简化版杀手游戏量化大模型的欺骗、侦查与判断能力
Deceive, Detect, and Disclose: Large Language Models Play Mini-Mafia
- 构建四人简化杀手游戏,通过数学公式预测黑帮胜率
- 实测显示模型间侦查与欺骗能力差异显著,如Grok 3 Mini最擅侦测
- 提供可量化的基准,适合评估大模型社交智能与协作潜力
大型语言模型在多智能体场景中的表现日益重要,但现有研究多为经验性,缺乏理论理解。为此,我们提出「Mini-Mafia」——一个四人简化版社会推理游戏,固定夜晚阶段使博弈聚焦于黑帮、侦探与村民间的单一关键互动。我们发现黑帮胜率 $p$ 可由公式 $ ext{logit}(p) = v imes (m - d)$ 精确预测,其中 $m$、$d$、$v$ 分别代表黑帮的欺骗能力、侦探的披露能力与村民的检测能力。基于此,我们构建了「Mini-Mafia基准」,利用贝叶斯推断从对战数据中估计各模型的内在参数 $m$、$d$、$v$。对于 $I$ 个模型,仅需 $3I$ 个参数即可预测全部 $I^3$ 种对局结果;五折交叉验证下,该公式相较随机基线提升 76.6% 的 Brier 得分。基准还揭示反直觉现象:Grok 3 Mini 是最强探测者,GPT-5 Mini 是最强披露者,均优于 DeepSeek V3.1、Claude Opus 4 和 Claude Sonnet 4;而 Claude Sonnet 4 检测能力接近随机水平。结果表明,尽管结构简单,Mini-Mafia 具有解析描述能力,是评估语言模型交互行为的可靠基准。
原文摘要 · Abstract (English)
Large language models are increasingly deployed in multi-agent settings whose outcomes hinge on social intelligence, motivating evaluations of their interactive capabilities; yet existing studies remain overwhelmingly empirical, leaving us without a theoretical understanding of how agent interactions determine collective outcomes. To address this, we introduce \textit{Mini-Mafia}, a four-player simplification of the social deduction game Mafia in which a fixed night phase reduces the game to a single critical exchange among a mafioso, a detective, and a villager. In this setting, we show that the mafia win-rate $p$ is predicted by the analytical formula $\text{logit}(p) = v \times (m - d)$, where $m$, $d$, and $v$ represent the mafioso's deception, the detective's disclosure, and the villager's detection capabilities. We turn this analytical framework into the \textit{Mini-Mafia Benchmark}, where Bayesian inference over gameplay data yields per-model estimates of the intrinsic parameters $m$, $d$, and $v$. For $I$ models, only $3I$ parameters suffice to predict the outcomes of all $I^3$ tournament combinations; and in 5-fold cross-validation the formula achieves a $76.6\%$ Brier-score reduction over a random baseline. The benchmark also reveals counterintuitive results: Grok 3 Mini is the strongest detector and GPT-5 Mini the strongest discloser, both ahead of DeepSeek V3.1, Claude Opus 4, and Claude Sonnet 4; while Claude Sonnet 4 is the weakest detector, near random chance. Together, these results show that Mini-Mafia, a simple but nontrivial multi-agent system, admits an analytical description and serves as a principled benchmark for language model interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。