用小贴纸干扰机器人视觉动作模型的去噪路径,让其完全失效
DRIFT: Derailing Denoising Trajectories of Flow-Matching VLAs with Adversarial Patch Attack

- 在输入端投放通用对抗贴片,破坏去噪轨迹的第一步
- 仅攻击首步即能破解全部任务,效果远超传统方法
- 适用于测试时攻击现成机器人策略,极具隐蔽性
基于流匹配的视觉-语言-动作(VLA)模型如pi0通过学习去噪速度场生成机器人动作,据称对对抗扰动具有鲁棒性。我们发现这种鲁棒性很大程度上是虚假的:源于以往攻击忽略了多步去噪微分方程。本文提出DRIFT(通过输入扰动引导去噪轨迹),在机器人夹持器上放置一个测试时通用对抗贴片,直接攻击现成策略的去噪速度场。核心发现反直觉:仅攻击第一去噪步骤比攻击更宽窗口的步骤更强且成本更低,这由输入空间优化中的梯度冲突导致,与训练时后门攻击情形恰好相反。在四个LIBERO套件中,DRIFT以单个小贴片成功破坏了所有原本可解的任务,显著超越动作空间和嵌入空间攻击基线。
原文摘要 · Abstract (English)
Flow-matching vision-language-action (VLA) models such as pi0 generate robot actions by integrating a learned denoising velocity field, and have been reported to resist adversarial perturbations that readily fool autoregressive VLAs. We show that this robustness is largely illusory: it stems from prior attacks ignoring the multi-step denoising ODE. We introduce DRIFT (Denoising Redirection via Input perturbation of the Flow-matching Trajectory), a test-time universal adversarial patch placed on the robot's gripper that attacks the denoising velocity field of an off-the-shelf policy. Our central finding is counterintuitive: attacking only the first denoising step is both stronger and cheaper than attacking a wider window of steps, which we explain through a gradient conflict unique to input-space optimization and which is exactly opposite to the training-time backdoor regime. On pi0 and pi0.5 across four LIBERO suites, DRIFT breaks essentially all originally-solvable tasks with a small single patch, far exceeding action- and embedding-space attack baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。