arXiv:2601.04266cs.CRcs.LG2026-01被引 10

用机械臂初始状态作隐蔽后门,攻击视觉语言动作模型

State Backdoor: Towards Stealthy Real-world Poisoning Attack on Vision-Language-Action Model in State Space

  • 以机械臂初始状态为触发器,实现隐蔽攻击
  • 在5个真实任务中攻击成功率超90%,正常性能不受影响
  • 适合研究机器人安全或防御的学者关注

视觉-语言-动作(VLA)模型广泛应用于机器人等安全关键型具身智能系统。然而其复杂的多模态交互也带来了新的安全漏洞。本文研究了一种针对VLA模型的后门威胁:恶意输入引发目标错误行为,同时保持对干净数据的正常性能。现有后门方法主要依赖在视觉模态中插入可见触发器,但在真实环境中因环境变化导致鲁棒性差、隐蔽性不足。为此,我们提出State Backdoor,一种新颖且实用的后门攻击方法,利用机械臂的初始状态作为触发器。为优化触发器的隐蔽性和有效性,设计了基于偏好引导的遗传算法(PGA),高效搜索状态空间以找到最小但最有效的触发器。在五个代表性VLA模型和五个真实世界任务上的大量实验表明,该方法攻击成功率超过90%,且不影响良性任务性能,揭示了具身AI系统中一个未被充分探索的漏洞。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models are widely deployed in safety-critical embodied AI applications such as robotics. However, their complex multimodal interactions also expose new security vulnerabilities. In this paper, we investigate a backdoor threat in VLA models, where malicious inputs cause targeted misbehavior while preserving performance on clean data. Existing backdoor methods predominantly rely on inserting visible triggers into visual modality, which suffer from poor robustness and low insusceptibility in real-world settings due to environmental variability. To overcome these limitations, we introduce the State Backdoor, a novel and practical backdoor attack that leverages the robot arm's initial state as the trigger. To optimize trigger for insusceptibility and effectiveness, we design a Preference-guided Genetic Algorithm (PGA) that efficiently searches the state space for minimal yet potent triggers. Extensive experiments on five representative VLA models and five real-world tasks show that our method achieves over 90% attack success rate without affecting benign task performance, revealing an underexplored vulnerability in embodied AI systems.

后门攻击具身智能机器人安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。