LLM代理在多轮游戏中形成声誉与欺骗策略,揭示了记忆驱动的社会动态。
Trust, Lies, and Long Memories: Emergent Social Dynamics and Reputation in Multi-Round Avalon with LLM Agents

- 代理通过记忆过往行为建立声誉,影响团队信任度
- 高推理力下恶方早期放水比例达75%,显著高于低推理力时的36%
- 适合研究智能体社会行为与信任机制的学者参考
我们研究了大型语言模型代理在《抵抗:阿瓦隆》这一隐藏身份欺骗类游戏中产生的涌现社会动态。与以往关注单局表现的研究不同,本实验让代理在多轮对局中保留过往互动记忆,包括角色身份与行为表现,从而观察社会关系如何演化。在188场游戏中,两个关键现象浮现:第一,跨局记忆催生自然形成的声誉系统——代理会提及过去经历如“我警惕重复上局因过早信任导致的失误”;声誉具有角色依赖性:同一代理在正义阵营被评价为“坦率”,在邪恶阵营则被视为“隐秘”,高声誉者获得46%更高的团队邀请率;第二,更高推理能力支持更策略性的欺骗:邪恶方在高推理场景中更常主动放弃早期任务以积累信任,其成功率从低推理时的36%上升至75%。结果表明,多轮记忆交互可催生可量化的声誉与欺骗行为模式。
原文摘要 · Abstract (English)
We study emergent social dynamics in LLM agents playing The Resistance: Avalon, a hidden-role deception game. Unlike prior work on single-game performance, our agents play repeated games while retaining memory of previous interactions, including who played which roles and how they behaved, enabling us to study how social dynamics evolve. Across 188 games, two key phenomena emerge. First, reputation dynamics emerge organically when agents retain cross-game memory: agents reference past behavior in statements like "I am wary of repeating last game's mistake of over-trusting early success." These reputations are role-conditional: the same agent is described as "straightforward" when playing good but "subtle" when playing evil, and high-reputation players receive 46% more team inclusions. Second, higher reasoning effort supports more strategic deception: evil players more often pass early missions to build trust before sabotaging later ones, 75% in high-effort games vs 36% in low-effort games. Together, these findings show that repeated interaction with memory gives rise to measurable reputation and deception dynamics among LLM agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。