arXiv:2608.10393cs.AIcs.RO2026-08

用扩散模型生成自然的对抗补丁,骗机器人做错动作

Hidden in Plain Sight: Diffusion-Based Unrestricted Robotic Attacks on Vision-Language-Action Models

论文配图:Hidden in Plain Sight: Diffusion-Based Unrestricted Robotic Attacks on Vision-Language-Action Models
图 1 · 摘自论文原文
  • 用预训练扩散模型的潜在轨迹生成视觉自然的对抗补丁
  • 黑盒攻击仅需目标模型输出动作,成功率超现有方法
  • 在仿真和真实机器人上均有效,暴露VLA模型安全风险

视觉-语言-动作(VLA)模型在多样化的操作任务中展现出强大的机器人控制能力。然而,其对抗鲁棒性尚未得到充分探索,利用这一弱点可能造成物理世界危害。现有攻击多依赖像素空间扰动或白盒访问,导致明显伪影且难以在真实机器人系统中部署。本文提出DURA,一种基于扩散模型的无限制机器人攻击方法,可生成对视觉自然的对抗补丁。DURA支持白盒与黑盒攻击设置,其中黑盒场景仅需目标模型的预测动作。通过沿预训练扩散模型的潜在轨迹优化,DURA在保持补丁自然的同时,引导机器人执行攻击者指定的目标动作。大量仿真与真实世界实验表明,DURA持续优于现有方法。研究揭示了物理部署的VLA模型存在安全风险,亟需更强防御措施。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have shown strong capabilities in controlling robots across diverse manipulation tasks. However, their adversarial robustness remains largely underexplored, and exploiting this weakness can lead to physical-world harm. Existing attacks on VLA models often rely on pixel-space perturbations or white-box access, resulting in noticeable artifacts and limited deployability in real-world robotic systems. In this work, we propose DURA, a diffusion-based unrestricted robotic attack that generates visually natural adversarial patches for VLA models. DURA supports both white-box and black-box attack settings, where the black-box setting requires only the predicted actions of the victim model. By optimizing along the latent trajectory of a pretrained diffusion model, DURA generates visually natural patches while steering the robot toward attacker-specified target actions. Extensive experiments in both simulation and the real physical world show that DURA consistently outperforms existing methods. Our findings expose a safety risk for physically deployed VLA models and call for stronger defenses.

对抗攻击机器人安全扩散模型VLA模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。