发现视觉语言动作模型被攻击后会完全停止执行指令
FreezeVLA: Action-Freezing Attacks against Vision-Language-Action Models
- 通过最小最大双层优化生成能让机器人停机的对抗图像
- 攻击成功率平均达76.2%,且跨任务、跨提示具有强迁移性
- 揭示了机器人智能系统的关键安全漏洞,适合关注机器人安全的研究者
视觉语言动作(VLA)模型正推动机器人技术快速发展,使智能体能够理解多模态输入并执行复杂长时序任务。然而,其在对抗攻击下的安全性与鲁棒性仍缺乏深入研究。本文识别并形式化了一种关键脆弱性:对抗图像可使VLA模型‘冻结’,导致其忽略后续指令,从而在关键时刻造成机器人静止。为系统研究此威胁,我们提出新攻击框架FreezeVLA,采用最小最大双层优化生成并评估行动冻结攻击。在三个先进VLA模型和四个机器人基准上的实验表明,FreezeVLA平均攻击成功率达76.2%,显著优于现有方法。此外,生成的对抗图像表现出强迁移能力,单一图像即可在多种语言提示下引发系统瘫痪。研究揭示了VLA模型的重大安全隐患,亟需构建有效防御机制。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models are driving rapid progress in robotics by enabling agents to interpret multimodal inputs and execute complex, long-horizon tasks. However, their safety and robustness against adversarial attacks remain largely underexplored. In this work, we identify and formalize a critical adversarial vulnerability in which adversarial images can "freeze" VLA models and cause them to ignore subsequent instructions. This threat effectively disconnects the robot's digital mind from its physical actions, potentially inducing inaction during critical interventions. To systematically study this vulnerability, we propose FreezeVLA, a novel attack framework that generates and evaluates action-freezing attacks via min-max bi-level optimization. Experiments on three state-of-the-art VLA models and four robotic benchmarks show that FreezeVLA attains an average attack success rate of 76.2%, significantly outperforming existing methods. Moreover, adversarial images generated by FreezeVLA exhibit strong transferability, with a single image reliably inducing paralysis across diverse language prompts. Our findings expose a critical safety risk in VLA models and highlight the urgent need for robust defense mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。