首个针对视觉语言动作模型的物理安全红队框架,能有效触发并检测潜在危害行为。
RedVLA: Physical Red Teaming for Vision-Language-Action Models

- 通过构建风险场景与迭代优化,系统性诱发VLA模型的异常行为。
- 在6个主流VLA模型上实现最高95.5%的攻击成功率达,10次内完成优化。
- 适合关注AI物理安全、部署前风险评估的研究者与工程师使用。
视觉语言动作(VLA)模型的实际应用受限于不可预测且不可逆的物理伤害风险。然而,当前缺乏有效的机制在部署前主动发现此类安全风险。为此,我们提出首个面向VLA模型物理安全的红队框架RedVLA。该框架采用两阶段流程:(I)风险场景合成,从正常轨迹中识别关键交互区域,并将风险因子置于其中,以干扰VLA执行流程并诱导目标不安全行为;(II)风险放大,通过无梯度优化迭代调整风险因子状态,基于轨迹特征确保在异构模型间稳定诱发。在6个代表性VLA模型上的实验表明,RedVLA可揭示多样化的不安全行为,且在10次优化迭代内达到最高95.5%的攻击成功率(ASR)。为进一步缓解风险,我们提出SimpleVLA-Guard,一种基于RedVLA生成数据的轻量级安全防护机制。相关数据、资源与代码已公开。
原文摘要 · Abstract (English)
The real-world deployment of Vision-Language-Action (VLA) models remains limited by the risk of unpredictable and irreversible physical harm. However, we currently lack effective mechanisms to proactively detect these physical safety risks before deployment. To address this gap, we propose \textbf{RedVLA}, the first red teaming framework for physical safety in VLA models. We systematically uncover unsafe behaviors through a two-stage process: (I) \textbf{Risk Scenario Synthesis} constructs a valid and task-feasible initial risk scene. Specifically, it identifies critical interaction regions from benign trajectories and positions the risk factor within these regions, aiming to entangle it with the VLA's execution flow and elicit a target unsafe behavior. (II) \textbf{Risk Amplification} ensures stable elicitation across heterogeneous models. It iteratively refines the risk factor state through gradient-free optimization guided by trajectory features. Experiments on six representative VLA models show that RedVLA uncovers diverse unsafe behaviors and achieves the ASR up to 95.5\% within 10 optimization iterations. To mitigate these risks, we further propose SimpleVLA-Guard, a lightweight safety guard built from RedVLA-generated data. Our data, assets, and code are available \href{https://redvla.github.io}{here}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。