arXiv:2511.12149cs.CRcs.AI2025-11被引 12

首个统一评估视觉语言动作模型安全性的框架,揭示其易受精准长序列攻击。

AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models

  • 构建统一评估框架,支持仿真与真实机器人场景测试
  • 发现现有攻击多导致无目标失败,缺乏精准长序列操控能力
  • 提出BackdoorVLA攻击,实测成功率最高达100%

视觉语言动作(VLA)模型使机器人能够理解自然语言指令并执行多样化任务,但其感知、语言与控制的融合带来了新的安全风险。尽管攻击研究日益增多,现有方法的有效性仍不明确,主要因缺乏统一评估框架。一个关键问题是不同VLA架构使用各异的动作分词器,影响复现性与公平比较。更重要的是,多数攻击未在真实场景中验证。为此,我们提出AttackVLA,一个契合VLA开发全周期的统一框架,涵盖数据构建、模型训练与推理阶段。在此框架下,我们实现多种攻击,包括所有现有针对VLA的攻击及适配自视觉语言模型的攻击,并在仿真与真实机器人环境中进行评估。分析表明:当前方法多引发非目标失效或静态动作状态,对驱动模型执行指定长序列动作的目标攻击研究严重不足。为此,我们提出BackdoorVLA,一种目标型后门攻击,当触发条件出现时,可迫使VLA执行攻击者指定的长序列动作。该攻击在模拟基准和真实机器人设置中均被验证,平均目标成功率58.4%,部分任务达100%。本工作为评估VLA漏洞提供了标准化工具,揭示了精确对抗操纵的可能性,推动对基于体化系统的安全性研究。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models enable robots to interpret natural-language instructions and perform diverse tasks, yet their integration of perception, language, and control introduces new safety vulnerabilities. Despite growing interest in attacking such models, the effectiveness of existing techniques remains unclear due to the absence of a unified evaluation framework. One major issue is that differences in action tokenizers across VLA architectures hinder reproducibility and fair comparison. More importantly, most existing attacks have not been validated in real-world scenarios. To address these challenges, we propose AttackVLA, a unified framework that aligns with the VLA development lifecycle, covering data construction, model training, and inference. Within this framework, we implement a broad suite of attacks, including all existing attacks targeting VLAs and multiple adapted attacks originally developed for vision-language models, and evaluate them in both simulation and real-world settings. Our analysis of existing attacks reveals a critical gap: current methods tend to induce untargeted failures or static action states, leaving targeted attacks that drive VLAs to perform precise long-horizon action sequences largely unexplored. To fill this gap, we introduce BackdoorVLA, a targeted backdoor attack that compels a VLA to execute an attacker-specified long-horizon action sequence whenever a trigger is present. We evaluate BackdoorVLA in both simulated benchmarks and real-world robotic settings, achieving an average targeted success rate of 58.4% and reaching 100% on selected tasks. Our work provides a standardized framework for evaluating VLA vulnerabilities and demonstrates the potential for precise adversarial manipulation, motivating further research on securing VLA-based embodied systems.

VLA安全后门攻击机器人对抗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。