提出部分可观测攻击框架,让对抗补丁在机器人长时任务中持续失效。
Partially Observable Adversarial Patch Attacks on Vision-Language-Action Models in Robotics

- 仅利用轨迹前缀定位关键视觉区域,生成固定补丁。
- 使目标物体语义错位并增加动作轨迹曲率,双重破坏感知与控制。
- 可在真实机器人上引发长时间任务失败,适合安全评估使用。
视觉-语言-动作(VLA)模型在机器人领域日益受到关注,但其对对抗攻击的鲁棒性仍缺乏研究。现有工作表明对抗补丁可误导基于VLA的机器人,但假设攻击者能完全访问整个执行轨迹,这在实践中不现实。本文提出一种部分可观测威胁模型:攻击者仅能利用轨迹的短前缀,生成一个固定的补丁应用于后续所有帧。为此,我们设计两阶段框架:首先通过模型注意力图定位与完整指令对应的视觉关键区域;随后优化补丁,破坏目标物体的语义定位,并增加动作轨迹的曲率,从而在感知和控制层面叠加失效。在仿真和真实机器人环境中的大量实验表明,该方法在部分可观测条件下仍能维持对抗效果,导致长时程任务中断,显著降低任务成功率。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models are gaining attention in robotics, yet their robustness to adversarial attacks remains largely unexplored. Existing work shows that adversarial patches can mislead VLA-based robots but assumes full access to the entire execution trajectory, an unrealistic requirement in practice. We address this limitation by formulating a partially observable threat model, where the adversary can exploit only a short prefix of the trajectory to generate a fixed patch applied to all subsequent frames. Under this setting, we propose a two-phase framework. First, we localize the patch using the model's attention maps to identify visually critical regions that correspond to the full instruction. Then, we optimize the patch to disrupt the semantic grounding of target objects and increase the curvature of action trajectories, thereby compounding failures in both perception and control. Extensive experiments in simulation and real-world robotic environments show that our method sustains adversarial effects under partial observability, inducing long-horizon disruptions and significantly reducing task success rates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。