arXiv:2606.27909cs.CLcs.AI2026-06

引入小丑角色测试大模型多跳心智推理能力。

Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs

论文配图:Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs
图 1 · 摘自论文原文
  • 设计三边博弈结构,小丑胜率靠被投票出局
  • GPT-4.1等模型误投小丑,自损己方利益
  • 仅DeepSeek掌握隐蔽可疑策略,适合研究多智能体推理

大型语言模型的心智推理评估通常基于双人社交推断游戏,其中每个线索只指向一个隐藏立场,导致模型仅依赖语言先验即可高分,无需模拟对手动机。本文将狼人杀游戏扩展为三方博弈,加入小丑角色——其获胜条件是被投票出局,因此与他人利益相反。在60场游戏中,小丑胜率高达60%-70%,而狼人从未超过20%;GPT-4.1在60%-70%的对局中第一天就投票淘汰小丑,这是明显自损的行为。开启自学习后,DeepSeek和Llama表现提升,但GPT-4.1反而变差,代价由村民承担而非狼人。唯有DeepSeek学会在不显刻意的前提下制造可疑感,并从中获益最大。三边激励结构揭示了双人推断游戏所无法暴露的多智能体推理层面。

原文摘要 · Abstract (English)

Theory-of-mind evaluations of large language models typically use dyadic social-deduction games, where every observable cue points to a single hidden side, so a model with strong language priors can score well without ever simulating opponents' incentives. We extend the Werewolf game with a Jester, a third faction whose utility on peer suspicion is inverted because it wins by being voted out, so optimal play requires reasoning across three opposing utility functions. Across 60 games on GPT-4.1, DeepSeek-V3.1, and Llama-3.3-70B with Jester self-learning on and off, the Jester wins 60-70% of games while Werewolves never exceed 20%, and GPT-4.1 wolves vote the Jester out on day 1 in 60-70% of games, a strictly self-defeating action. Self-learning helps DeepSeek and Llama but hurts GPT-4.1, with the cost landing on Villagers rather than Werewolves. Only DeepSeek learns the subtle strategy of looking suspicious without looking intentionally suspicious, and it gains the most from the loop. Triadic incentive structure exposes a layer of multi-agent reasoning that dyadic deduction games leave invisible.

心智推理多智能体游戏博弈LLM评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。