通过游戏实测大模型的社会推理能力,区分其判断失误与行动失败。
MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games

- 在每轮游戏中秘密提问,探测模型真实信念而不影响游戏进程。
- 发现模型自信度严重失准,误判怀疑频率高出1.5倍,且多数决策难以通过重演改变。
- 适合研究大模型社会认知、信任机制或博弈行为的学者使用。
大模型在社交推理任务中的公开行为无法反映其真实心理状态:正确投票可能只是猜测,高超谎言也无从揭示其实际信念。我们提出MafiaScope,一个开放测试平台,将社交推理游戏Mafia转化为测量机器心智理论的工具。该系统能区分代理因误解局势而失败,还是虽有正确认知却未采取行动——这一差异仅靠结果和对话无法察觉。每轮公开发言后,各代理需私下回答结构化探针问题,回答不参与游戏且由引擎对照真实情况评分。交互式可视化器可从单个代理视角回放游戏,展示时间对齐的准确率与校准度,并支持任意步骤的反事实重演。在涵盖两组模型、数万条探针响应的案例研究中,我们发现模型的自信程度严重失准,其误判被怀疑频率比实际高出1.5倍;单次投票的反事实重演极少改变游戏结局:结果反转主要发生在代理已形成正确信念的情况下,而基于错误世界模型的决策在重演中基本保持不变。该平台、可视化工具、录制游戏及反事实重演数据集均以开源许可发布。代码:https://github.com/karpovilia/mafiascope。在线演示:https://karpovilia.github.io/mafiascope/。演示视频:https://vimeo.com/1208920221。
原文摘要 · Abstract (English)
An LLM agent's public behaviour reveals little about its social reasoning: an agent that votes correctly may be guessing, and an agent that lies well leaves no trace of what it actually believes. We present MafiaScope, an open testbed that turns the social deduction game Mafia into a measurement instrument for machine Theory of Mind. It distinguishes whether an agent lost because it misread the game or because it failed to act on a correct assessment, a distinction that is invisible from outcomes and dialogue transcripts alone. After every public utterance, each agent privately answers structured probe questions whose responses never re-enter the game and are scored against the ground truth known to the engine. An interactive visualizer replays games from the perspective of an individual agent's beliefs, displays timeline-aligned accuracy and calibration, and supports counterfactual replay from any recorded step. In a case study across two model families comprising tens of thousands of parsed probe responses, we find that stated confidence is poorly calibrated, agents overestimate how often they are suspected by a factor of 1.5, and single-vote counterfactual replays rarely change game outcomes: outcome flips occur primarily when the agent had already formed a correct belief state, whereas decisions made under an incorrect model of the world remain largely unchanged under resampling. The engine, visualizer, recorded games, and counterfactual replay corpus are released under an open-source licence. Code: https://github.com/karpovilia/mafiascope. Live demo: https://karpovilia.github.io/mafiascope/. Screencast: https://vimeo.com/1208920221.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。