通过可审计框架揭示大模型在隐藏信息博弈中的决策机制。
Auditing Belief-Conditioned LLM Agents in Hidden-Information Social Deduction Games
- 构建外部信念状态系统,记录信念更新与行动偏差作为证据
- 实验证明主动信念机制使好方胜率从20.5%提升至39.0%
- 适合研究多智能体博弈中模型行为可解释性与安全迭代
在9人狼人杀环境中,智能体在严格信息隔离下运行,我们构建了一个可审计框架,维护对隐藏身份的外部信念状态,记录信念更新与信念-行动偏差作为结构化证据,并支持离线回溯改进。在1080场冻结游戏中,包括信念禁用、主动信念、核心消融、阵营限制、消费策略和高负载等设置,以及200组种子配对的A0/A1对比,主动信念条件下好方胜率从0.205升至0.390(配对McNemar检验χ²=16.4,p<0.001),且减少不可逆的女巫毒杀错误。然而,该提升并非源于信念内容本身——直接行动-信念一致性仅约0.21,且仅向狼人提供信念反而比仅向好人提供更有利,反驳了简单持有者受益假说。因此,结论为关联性而非因果性。本研究贡献在于框架本身:它使效果可测量,暴露低一致性,拒绝不可靠干预,分离策略效应与负载混杂。我们主张,在高噪声隐藏信息游戏中,外部信念主要作为可审计认知基线,兼具决策信号,将黑箱行为转化为可重放证据,实现更安全可控的迭代。
原文摘要 · Abstract (English)
Evaluating LLM agents in hidden-information multi-agent settings is hard: final outcomes are high-variance and rarely reveal why an agent decided as it did. We study this in a 9-player Werewolf environment where agents act under strict, code-level information isolation, and we build an auditable framework that maintains an external belief state over hidden roles, logs belief updates and belief-action deviations as structured evidence, and supports a defensive offline improvement loop that reviews bad cases before any strategy change. Across 1,080 frozen games spanning belief-disabled, active-belief, kernel-ablation, camp-restricted, consumption-policy, and high-load arms, and including a seed-paired A0/A1 comparison, the active-belief condition is associated with substantially better good-side outcomes: in the 200-seed A0/A1 comparison the good-side win rate rises from 0.205 to 0.390 (paired McNemar $χ^2 = 16.4$, $p < 0.001$), with fewer irreversible witch-poison errors. We do not, however, attribute this shift to belief content. Direct action-belief consistency is low ($\approx 0.21$), and giving belief only to the werewolves helps the good side more than giving it only to the good side, which argues against a simple holder-benefit account; we therefore report the effect as an association and treat its mechanism as unresolved. The contribution is the audit framework itself: it makes the effect measurable, exposes low direct action-belief consistency, rejects an unreliable forced-consumption intervention with evidence, and separates strategy effects from load confounds. We accordingly position external belief in high-noise hidden-information games primarily as an auditable cognitive baseline that also carries decision-relevant signal, turning opaque agent behavior into replayable evidence for safer, controlled iteration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。