首个融合视觉与因果推理的社交推理游戏AI,能看懂动作猜身份。
CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games

- 通过视频分析玩家行为,用因果推理链推断隐藏身份
- 在人类参与测试中表现优于纯文本模型,互动更自然
- 适合研究多模态社交智能、人机协作的学者和开发者
社交推理游戏(如狼人杀)已成为检验人工智能复杂社交能力的挑战性场景。这类游戏需要推理、欺骗与合作等高阶社交技能。尽管大语言模型(LLMs)推动了游戏智能体的发展,但现有方法多为纯文本输入,忽视了人类社交中至关重要的多模态特性。为此,我们提出CaM-Wolf,首个集成多模态感知与生成的社交推理游戏智能体。该系统处理其他玩家的视频输入,采用基于强化学习训练的因果感知推理器,建立可观测行为与隐藏角色之间的逻辑关联,并通过动画化身进行自我呈现。实验与用户研究表明,CaM-Wolf在代理游戏表现上优于现有方法,显著提升了人机交互质量。本工作标志着向具备细腻社交能力的人类级智能体迈进的重要一步。代码已公开于https://3dagentworld.github.io/avatar_wolf。
原文摘要 · Abstract (English)
Social deduction games (SDGs) such as Werewolf have become challenging testbeds for AI agents. These games require complex social skills such as reasoning, deception, and collaboration. While recent advances in large language models (LLMs) have driven significant progress in SDG agents, current approaches are predominantly text-based, overlooking the multimodal nature that is fundamental to human social interaction. To bridge this gap, we introduce CaM-Wolf, the first SDG agent that integrates multimodal perception and generation. CaM-Wolf processes video inputs from other players, employs a causal-aware Reasoner trained via reinforcement learning to establish logical chains between observable behaviors and hidden roles, and presents itself through an animated avatar. Our experiments and user study show that CaM-Wolf achieves superior agent gameplay performance and enhances the quality of human-AI interaction. This work represents a significant advancement towards creating more human-like AI agents capable of participating in nuanced social dynamics. Our code is available at https://3dagentworld.github.io/avatar_wolf.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。