让视觉语言模型在3D游戏中通过言行双重欺骗来伪装身份,揭示非语言行为的关键作用。
Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions

- 构建3D多模态游戏环境,让智能体通过语言与动作联合欺骗对手。
- 非语言行为比语言更关键,是决定胜负的核心因素。
- 提供可配置框架与标注系统,适合安全对齐与认知研究。
大语言模型与视觉语言模型在战略欺骗中的表现已成为人工智能对齐与安全的核心关切。社交推理游戏(玩家持有隐藏身份并通过交流推断彼此角色)是典型测试场景,尤其在多智能体设置中。然而现有测试环境均为纯文本且固定配置,忽略了欺骗分类学中视为核心的非语言感官运动通道,导致观察到的行为究竟是模型本质还是外部框架所致难以判断。本文提出MineAmongUs——一个3D多模态的Among Us沙盒环境,让伪装者智能体需通过言语与非言语行为联合欺骗船员。同时提出ARIA,一种可配置的视觉语言模型智能体框架,暴露五个认知组件的消融轴;并设计原子级与弧级标注方案,基于欺骗分类学,由大模型作为裁判实现接近人类水平的标注一致性。实证结果表明,视觉语言模型智能体通过言语与非言语协同欺骗赢得比赛,且在框架消融与跨模型评估中,非语言通道均成为更具决定性的获胜因素。本工作为具身视觉语言模型对齐研究开辟了新路径。
原文摘要 · Abstract (English)
Strategic deception by LLM and VLM agents has emerged as a central AI alignment and safety concern. Social-deduction games (where each player holds a hidden role and communicates with others to deduce identities) serve as the canonical testbed, particularly in multi-agent settings. Existing testbeds, however, are text-only and run on a single fixed agent configuration, missing the non-verbal sensorimotor channels treated as core by deception taxonomies and leaving it ambiguous whether an observed behavior reflects the underlying model or the surrounding harness. We introduce MineAmongUs, a 3D multimodal Among Us sandbox where imposter agents must deceive crewmates through joint verbal and non-verbal action. We also propose ARIA, a configurable VLM-agent harness that exposes five cognitive-component ablation axes; and an atom- and arc-level annotation scheme grounded in deception taxonomies and operationalized at scale by an LLM-as-a-Judge reaching near-human atom-labeling agreement. Empirical results show that VLM agents pursue imposter wins through joint verbal and non-verbal deception, with non-verbal channels emerging as the more decisive winning contributors across both harness ablation and cross-VLM evaluation. Taken together, our work opens a new path for embodied VLM-agent alignment research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。