arXiv:2511.16203cs.CVcs.AI2025-11被引 17

研究视觉语言动作模型在多模态攻击下的脆弱性,揭示其决策易受干扰。

When Alignment Fails: Multimodal Adversarial Attacks on Vision-Language-Action Models

  • 设计三类多模态攻击:文本扰动、视觉干扰和跨模态语义错位。
  • 微小扰动即可导致机器人行为严重偏离,暴露对对齐的依赖。
  • 提出首个自动语义引导提示框架,适用于安全评估与鲁棒性研究。

视觉语言动作模型(VLAs)在具身环境中展现出卓越进展,使机器人能够通过统一的多模态理解实现感知、推理与行动。尽管能力突出,其对抗鲁棒性仍缺乏系统研究,尤其在真实多模态与黑盒条件下。现有工作多关注单模态扰动,忽视了影响具身推理的核心跨模态对齐问题。本文提出VLA-Fool,首次在白盒与黑盒设置下全面评估具身VLAs的多模态对抗鲁棒性。该方法涵盖三个层级攻击:(1) 基于梯度与提示的文本扰动,(2) 通过补丁与噪声的视觉干扰,(3) 故意破坏感知与指令间语义对应关系的跨模态错位攻击。进一步引入面向VLAs的语义空间,构建首个自动生成且语义引导的提示框架。在LIBERO基准上使用微调后的OpenVLA模型进行实验,结果表明微小多模态扰动即可引发显著行为偏差,凸显具身多模态对齐的脆弱性。

原文摘要 · Abstract (English)

Vision-Language-Action models (VLAs) have recently demonstrated remarkable progress in embodied environments, enabling robots to perceive, reason, and act through unified multimodal understanding. Despite their impressive capabilities, the adversarial robustness of these systems remains largely unexplored, especially under realistic multimodal and black-box conditions. Existing studies mainly focus on single-modality perturbations and overlook the cross-modal misalignment that fundamentally affects embodied reasoning and decision-making. In this paper, we introduce VLA-Fool, a comprehensive study of multimodal adversarial robustness in embodied VLA models under both white-box and black-box settings. VLA-Fool unifies three levels of multimodal adversarial attacks: (1) textual perturbations through gradient-based and prompt-based manipulations, (2) visual perturbations via patch and noise distortions, and (3) cross-modal misalignment attacks that intentionally disrupt the semantic correspondence between perception and instruction. We further incorporate a VLA-aware semantic space into linguistic prompts, developing the first automatically crafted and semantically guided prompting framework. Experiments on the LIBERO benchmark using a fine-tuned OpenVLA model reveal that even minor multimodal perturbations can cause significant behavioral deviations, demonstrating the fragility of embodied multimodal alignment.

多模态攻击具身智能鲁棒性对齐脆弱性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。