精准干扰视觉语言模型关键物体,让机器人误判却不失效。
ADVEDM:Fine-grained Adversarial Attack against VLM-based Embodied Agents
- 只修改图像中少数关键物体的语义,保持其他区域不变
- 在真实任务场景中成功诱导机器人做出错误但合理的动作决策
- 适合研究VLM安全性和智能体对抗攻击的开发者
视觉语言模型(VLM)凭借强大的推理与规划能力,被广泛应用于机器人等具身智能体的决策任务,如自动驾驶和机械臂操作。现有对抗攻击或依赖完全已知目标模型(不现实),或因破坏过多图像语义信息导致感知与任务上下文不一致,打断推理流程,产生无效输出,无法影响物理世界交互。为此,本文提出细粒度对抗攻击框架ADVEDM,仅修改少数关键物体的感知信息,同时保留其余区域语义。该方法有效降低感知与任务上下文的冲突,使VLM生成看似合理却错误的决策,进而影响智能体行为,构成更真实的物理世界安全威胁。设计了两种变体:ADVEDM-R移除特定物体语义,ADVEDM-A添加新物体语义。在通用场景与具身决策任务中的实验表明,该方法具备精细控制能力和卓越攻击效果。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs), with their strong reasoning and planning capabilities, are widely used in embodied decision-making (EDM) tasks in embodied agents, such as autonomous driving and robotic manipulation. Recent research has increasingly explored adversarial attacks on VLMs to reveal their vulnerabilities. However, these attacks either rely on overly strong assumptions, requiring full knowledge of the victim VLM, which is impractical for attacking VLM-based agents, or exhibit limited effectiveness. The latter stems from disrupting most semantic information in the image, which leads to a misalignment between the perception and the task context defined by system prompts. This inconsistency interrupts the VLM's reasoning process, resulting in invalid outputs that fail to affect interactions in the physical world. To this end, we propose a fine-grained adversarial attack framework, ADVEDM, which modifies the VLM's perception of only a few key objects while preserving the semantics of the remaining regions. This attack effectively reduces conflicts with the task context, making VLMs output valid but incorrect decisions and affecting the actions of agents, thus posing a more substantial safety threat in the physical world. We design two variants of based on this framework, ADVEDM-R and ADVEDM-A, which respectively remove the semantics of a specific object from the image and add the semantics of a new object into the image. The experimental results in both general scenarios and EDM tasks demonstrate fine-grained control and excellent attack performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。