用视觉触发让机器人在特定时刻执行指定动作,几乎不影响正常任务表现。
DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models

- 通过分段微调的窗口一致重标记法,实现对动作级别的隐蔽攻击。
- 仅用0.31%污染数据,攻击成功率高达98.67%~99.83%,且正常任务保留率超98.5%。
- 适用于实际机器人场景,对视角变化鲁棒,适合研究安全与对抗性风险者。
视觉-语言-动作(VLA)模型将多模态感知与语言指令映射为可执行机器人动作,易受行为后门攻击:训练中引入的隐藏触发器可在不破坏正常任务性能的前提下诱导非预期物理动作。现有工作多关注无目标攻击或任务级劫持,缺乏对单个动作的精细控制。本文提出DropVLA,一种在真实黑盒设置下、数据污染受限的行动级后门攻击方法,采用窗口一致重标记策略进行分块微调。在OpenVLA-7B上,仅使用0.31%污染片段,纯视觉中毒即可达到98.67%-99.83%攻击成功率(ASR),同时保持98.50%-99.17%的干净任务保留率,并在500 Hz(0.05秒内)25步控制内触发目标动作。纯文本触发在低污染预算下不稳定,结合文本与视觉未带来一致提升。后门对适度触发变化具有鲁棒性,跨评估套件迁移率分别为96.27%和99.09%,而纯文本攻击基本失效(0.72%)。进一步在7-DoF Franka机械臂上通过pi0-fast验证了现实世界可行性,证明在相机相对运动导致图像平面触发漂移时仍具显著攻击效果。结果表明,VLA模型可被极小污染量悄然操控至关键动作粒度,且不引发明显性能下降。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models map multimodal perception and language instructions to executable robot actions, making them particularly vulnerable to behavioral backdoor manipulation: a hidden trigger introduced during training can induce unintended physical actions while nominal task performance remains intact. Prior work on VLA backdoors primarily studies untargeted attacks or task-level hijacking, leaving fine-grained control over individual actions largely unexplored. In this work, we present DropVLA, an action-level backdoor attack that forces a reusable action primitive (e.g., open_gripper) to execute at attacker-chosen decision points under a realistic pipeline-black-box setting with limited data-poisoning access, using a window-consistent relabeling scheme for chunked fine-tuning. On OpenVLA-7B evaluated with LIBERO, vision-only poisoning achieves 98.67%-99.83% attack success rate (ASR) with only 0.31% poisoned episodes while preserving 98.50%-99.17% clean-task retention, and successfully triggers the targeted action within 25 control steps at 500 Hz (0.05 s). Text-only triggers are unstable at low poisoning budgets, and combining text with vision provides no consistent ASR improvement over vision-only attacks. The backdoor remains robust to moderate trigger variations and transfers across evaluation suites (96.27%, 99.09%), whereas text-only largely fails (0.72%). We further validate physical-world feasibility on a 7-DoF Franka arm with pi0-fast, demonstrating non-trivial attack efficacy under camera-relative motion that induces image-plane trigger drift. These results reveal that VLA models can be covertly steered at the granularity of safety-critical actions with minimal poisoning and without observable degradation of nominal performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。