arXiv:2506.03350cs.ROcs.AI2025-06被引 35

攻击视觉语言动作模型,用文本指令实现机器人完全控制

Adversarial Attacks on Robotic Vision Language Action Models

  • 将大模型越狱攻击方法移植到机器人系统,首次实现对VLA的完整控制
  • 单次文本攻击即可让机器人在长序列中持续达成目标动作空间
  • 适用于研究机器人安全或对抗攻防的科研人员与工程师

视觉语言动作模型(VLAs)通过融合多模态感知输入,在端到端控制中推动了机器人领域的发展,其能力主要源于基于前沿大语言模型(LLMs)的架构。然而,LLMs已知易受对抗性滥用,而机器人存在显著物理风险,因此亟需评估VLAs是否继承此类脆弱性。本文首次系统研究对VLA控制机器人的对抗攻击。核心贡献是将大语言模型越狱攻击方法适配并应用于真实机器人场景,发现仅需在任务初始阶段施加一次文本攻击,即可实现对常用VLAs动作空间的完全可达,且攻击效果在长时间轨迹中仍能维持。这一现象与传统越狱攻击不同,因现实世界中的攻击无需语义上关联伤害概念。相关代码已开源于https://github.com/eliotjones1/robogcg。

原文摘要 · Abstract (English)

The emergence of vision-language-action models (VLAs) for end-to-end control is reshaping the field of robotics by enabling the fusion of multimodal sensory inputs at the billion-parameter scale. The capabilities of VLAs stem primarily from their architectures, which are often based on frontier large language models (LLMs). However, LLMs are known to be susceptible to adversarial misuse, and given the significant physical risks inherent to robotics, questions remain regarding the extent to which VLAs inherit these vulnerabilities. Motivated by these concerns, in this work we initiate the study of adversarial attacks on VLA-controlled robots. Our main algorithmic contribution is the adaptation and application of LLM jailbreaking attacks to obtain complete control authority over VLAs. We find that textual attacks, which are applied once at the beginning of a rollout, facilitate full reachability of the action space of commonly used VLAs and often persist over longer horizons. This differs significantly from LLM jailbreaking literature, as attacks in the real world do not have to be semantically linked to notions of harm. We make all code available at https://github.com/eliotjones1/robogcg .

机器人安全对抗攻击视觉语言动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。