arXiv:2509.03383cs.AIcs.RO2025-09被引 2

首次系统研究机器人视觉语言动作模型的安全漏洞,揭示其物理攻击风险。

ANNIE: Be Careful of Your Robots

  • 构建基于物理约束的安全违规分类体系,区分严重、危险、高风险三类
  • 设计9个场景2400段视频-动作序列的基准测试集ANNIEBench
  • 提出任务感知对抗攻击框架,实测攻击成功率超50%,可远程操控机器人

将视觉-语言-动作(VLA)模型融入具身人工智能(EAI)机器人正迅速提升其在人类环境中的复杂长程任务能力。然而,这类系统引入了关键安全风险:被攻破的VLA模型可将感知输入的对抗扰动直接转化为不安全的物理动作。传统机器学习安全定义与方法已不再适用。具身智能系统带来新问题:何为安全?如何衡量?如何在物理交互环境中设计有效攻防机制?本文首次对具身智能系统的对抗安全攻击进行系统研究,依据人机交互国际标准建立安全违规分类体系(严重、危险、高风险),基于分离距离、速度、碰撞边界等物理约束进行形式化定义;提出ANNIEBench基准,包含9个安全关键场景,共2400段视频-动作序列,用于评估具身安全性能;设计ANNIE-Attack框架,通过攻击主控模型将长程目标分解为帧级扰动。在代表性EAI模型上的评估显示,所有安全类别攻击成功率均超过50%。进一步验证了稀疏且自适应的攻击策略,并通过实体机器人实验确认其现实影响。结果揭示具身智能系统中此前未受重视但后果严重的攻击面,凸显物理人工智能时代安全防御的紧迫性。代码开源于https://github.com/RLCLab/Annie。

原文摘要 · Abstract (English)

The integration of vision-language-action (VLA) models into embodied AI (EAI) robots is rapidly advancing their ability to perform complex, long-horizon tasks in humancentric environments. However, EAI systems introduce critical security risks: a compromised VLA model can directly translate adversarial perturbations on sensory input into unsafe physical actions. Traditional safety definitions and methodologies from the machine learning community are no longer sufficient. EAI systems raise new questions, such as what constitutes safety, how to measure it, and how to design effective attack and defense mechanisms in physically grounded, interactive settings. In this work, we present the first systematic study of adversarial safety attacks on embodied AI systems, grounded in ISO standards for human-robot interactions. We (1) formalize a principled taxonomy of safety violations (critical, dangerous, risky) based on physical constraints such as separation distance, velocity, and collision boundaries; (2) introduce ANNIEBench, a benchmark of nine safety-critical scenarios with 2,400 video-action sequences for evaluating embodied safety; and (3) ANNIE-Attack, a task-aware adversarial framework with an attack leader model that decomposes long-horizon goals into frame-level perturbations. Our evaluation across representative EAI models shows attack success rates exceeding 50% across all safety categories. We further demonstrate sparse and adaptive attack strategies and validate the real-world impact through physical robot experiments. These results expose a previously underexplored but highly consequential attack surface in embodied AI systems, highlighting the urgent need for security-driven defenses in the physical AI era. Code is available at https://github.com/RLCLab/Annie.

具身智能对抗攻击机器人安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。