用多样化对抗策略暴露视觉语言动作模型的语义脆弱性
Uncovering Linguistic Fragility in Vision-Language-Action Models via Diversity-Aware Red Teaming
- 设计多样性感知的红队框架,生成多样且有效的对抗指令
- 使主流模型任务成功率从93.33%降至5.85%,暴露严重安全盲点
- 适合评估机器人智能体安全性,尤其关注语义细微差异的影响
视觉语言动作(VLA)模型在机器人操作中取得显著进展,但其对语言细微差别的鲁棒性仍是关键且未被充分探索的安全隐患,威胁真实场景部署。红队测试通过识别引发灾难性行为的环境场景,是保障具身智能体安全的关键步骤。强化学习(RL)已成为自动化红队的有力方法,旨在发现这些漏洞。然而,标准的基于RL的对抗者常因奖励最大化导致严重模式崩溃,收敛至少数琐碎或重复的失败模式,无法揭示有意义风险的完整图景。为此,我们提出新型多样性感知具身红队(DAERT)框架,以暴露VLA对语言变化的脆弱性。该框架基于统一策略,能生成多样且具挑战性的指令,同时保证攻击有效性(以物理仿真中的执行失败为衡量标准)。我们在多个机器人基准上对两种先进VLA模型(π₀和OpenVLA)进行了广泛实验。结果表明,该方法持续发现更广泛、更有效的对抗指令,将平均任务成功率从93.33%降至5.85%,证明了其在规模化压力测试和暴露关键安全盲点方面的有效性。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have achieved remarkable success in robotic manipulation. However, their robustness to linguistic nuances remains a critical, under-explored safety concern, posing a significant safety risk to real-world deployment. Red teaming, or identifying environmental scenarios that elicit catastrophic behaviors, is an important step in ensuring the safe deployment of embodied AI agents. Reinforcement learning (RL) has emerged as a promising approach in automated red teaming that aims to uncover these vulnerabilities. However, standard RL-based adversaries often suffer from severe mode collapse due to their reward-maximizing nature, which tends to converge to a narrow set of trivial or repetitive failure patterns, failing to reveal the comprehensive landscape of meaningful risks. To bridge this gap, we propose a novel \textbf{D}iversity-\textbf{A}ware \textbf{E}mbodied \textbf{R}ed \textbf{T}eaming (\textbf{DAERT}) framework, to expose the vulnerabilities of VLAs against linguistic variations. Our design is based on evaluating a uniform policy, which is able to generate a diverse set of challenging instructions while ensuring its attack effectiveness, measured by execution failures in a physical simulator. We conduct extensive experiments across different robotic benchmarks against two state-of-the-art VLAs, including $π_0$ and OpenVLA. Our method consistently discovers a wider range of more effective adversarial instructions that reduce the average task success rate from 93.33\% to 5.85\%, demonstrating a scalable approach to stress-testing VLA agents and exposing critical safety blind spots before real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。