arXiv:2505.16640cs.CRcs.AI2025-05NeurIPS被引 46

首次揭示视觉语言动作模型的后门漏洞,可隐蔽触发错误行为。

BadVLA: Towards Backdoor Attacks on Vision-Language-Action Models via Objective-Decoupled Optimization

论文配图:BadVLA: Towards Backdoor Attacks on Vision-Language-Action Models via Objective-Decoupled Optimization
图 1 · 摘自论文原文
  • 分离特征空间并仅在触发器出现时改变控制输出
  • 攻击成功率接近100%,对正常任务影响极小
  • 适合关注机器人安全与可信智能系统的研究者

视觉语言动作(VLA)模型通过多模态输入实现端到端决策,推动了机器人控制的发展。然而其紧密耦合的架构带来了新的安全风险。不同于传统对抗扰动,后门攻击具有隐蔽性、持续性和实际威胁性,尤其在新兴的训练即服务模式下更为突出,但此前在VLA模型中尚未被深入研究。为此,我们提出基于目标解耦优化的BadVLA攻击方法,首次系统揭示了VLA模型的后门漏洞。该方法包含两阶段:(1) 显式分离特征空间,将触发器表示与正常输入分开;(2) 仅在触发器存在时产生条件性控制偏差,同时保持干净任务性能。多个VLA基准测试的实验结果表明,BadVLA在保持高清洁任务准确率的前提下,攻击成功率接近100%。进一步分析证实其对常见输入扰动、任务迁移和模型微调均具有鲁棒性,凸显当前VLA部署中的关键安全隐患。本工作为构建安全可信的具身智能模型提供了重要警示。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have advanced robotic control by enabling end-to-end decision-making directly from multimodal inputs. However, their tightly coupled architectures expose novel security vulnerabilities. Unlike traditional adversarial perturbations, backdoor attacks represent a stealthier, persistent, and practically significant threat-particularly under the emerging Training-as-a-Service paradigm-but remain largely unexplored in the context of VLA models. To address this gap, we propose BadVLA, a backdoor attack method based on Objective-Decoupled Optimization, which for the first time exposes the backdoor vulnerabilities of VLA models. Specifically, it consists of a two-stage process: (1) explicit feature-space separation to isolate trigger representations from benign inputs, and (2) conditional control deviations that activate only in the presence of the trigger, while preserving clean-task performance. Empirical results on multiple VLA benchmarks demonstrate that BadVLA consistently achieves near-100% attack success rates with minimal impact on clean task accuracy. Further analyses confirm its robustness against common input perturbations, task transfers, and model fine-tuning, underscoring critical security vulnerabilities in current VLA deployments. Our work offers the first systematic investigation of backdoor vulnerabilities in VLA models, highlighting an urgent need for secure and trustworthy embodied model design practices. We have released the project page at https://badvla-project.github.io/.

后门攻击机器人安全多模态模型可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。