arXiv:2604.23775cs.RO2026-04被引 11

VLA模型安全全景解析:从威胁到防御的系统性梳理

Vision-Language-Action Safety: Threats, Challenges, Evaluations, and Mechanisms

  • 按攻击/防御发生时间分轴,构建VLA安全框架
  • 揭示多模态攻击与长时轨迹误差传播等核心风险
  • 适合机器人、AI安全及具身智能研究者参考

视觉-语言-动作(VLA)模型正成为具身智能的统一基础,但其物理交互特性带来了新安全挑战:不可逆物理后果、跨视觉、语言与状态的多模态攻击面、防御实时性要求、长时轨迹中的错误传播,以及数据供应链漏洞。现有研究分散于机器人学习、对抗机器学习、AI对齐和自主系统安全领域。本文提供关于VLA模型安全的统一、最新综述。沿攻击与防御的时间轴(训练期与推理期)组织内容,将每类威胁与可缓解阶段对应。首先界定VLA安全范畴,区分于文本大模型与传统机器人安全,并回顾VLA模型架构、训练范式与推理机制。随后通过四维视角审视文献:攻击、防御、评估与部署。涵盖训练期威胁如数据投毒与后门,以及推理期攻击如对抗补丁、跨模态扰动、语义越狱和冻结攻击。综述训练期与运行时防御策略,分析现有基准与度量方法,并讨论六个部署场景中的安全挑战。最后提出关键开放问题,包括具身轨迹的可证明鲁棒性、物理可实现的防御、安全感知训练、统一运行时安全架构及标准化评估。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models are emerging as a unified substrate for embodied intelligence. This shift raises a new class of safety challenges, stemming from the embodied nature of VLA systems, including irreversible physical consequences, a multimodal attack surface across vision, language, and state, real-time latency constraints on defense, error propagation over long-horizon trajectories, and vulnerabilities in the data supply chain. Yet the literature remains fragmented across robotic learning, adversarial machine learning, AI alignment, and autonomous systems safety. This survey provides a unified and up-to-date overview of safety in Vision-Language-Action models. We organize the field along two parallel timing axes, attack timing (training-time vs. inference-time and defense timing (training-time vs. inference-time, linking each class of threat to the stage at which it can be mitigated. We first define the scope of VLA safety, distinguishing it from text-only LLM safety and classical robotic safety, and review the foundations of VLA models, including architectures, training paradigms, and inference mechanisms. We then examine the literature through four lenses: Attacks, Defenses, Evaluation, and Deployment. We survey training-time threats such as data poisoning and backdoors, as well as inference-time attacks including adversarial patches, cross-modal perturbations, semantic jailbreaks, and freezing attacks. We review training-time and runtime defenses, analyze existing benchmarks and metrics, and discuss safety challenges across six deployment domains. Finally, we highlight key open problems, including certified robustness for embodied trajectories, physically realizable defenses, safety-aware training, unified runtime safety architectures, and standardized evaluation.

VLA安全具身智能多模态攻击防御机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。