发现视觉语言动作模型在真实干扰下推理与驾驶轨迹易被攻击,威胁自动驾驶安全。
ReasonBreak: Probing Vulnerabilities in Reasoning-Enabled Vision-Language-Action Models for Autonomous Driving

- 针对文本输入扰动设计黑盒攻击,测试推理与驾驶行为的脆弱性。
- 在闭环仿真中攻击成功率高达89%(推理)和72%(轨迹),碰撞率上升。
- 提出包含语义与结构评估的安全评测框架,适合自动驾驶安全研究者使用。
具备推理能力的视觉语言动作(VLA)模型被用于端到端自动驾驶,假设推理与轨迹生成紧密耦合。然而,这类系统在真实输入扰动下的鲁棒性尚未得到充分探索。我们发现这些模型对现实中的输入扰动高度敏感,在闭环仿真中推理攻击成功率达89%,轨迹操纵成功率达72%,导致碰撞率上升和安全指标恶化。以NVIDIA最新Alpamayo模型为例,开展首个针对推理型VLA模型的系统性黑盒研究,评估文本输入污染对推理与驾驶行为的影响。提出一种兼顾语义与结构的推理感知评估框架,以及以安全为核心的度量方法,并构建了用于评估推理-轨迹交互攻击与防御的基准。结果表明,亟需更严格的评估机制与增强防御措施,以保障自动驾驶中推理型VLA系统的安全性。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models with integrated reasoning have been proposed for end-to-end autonomous driving, assuming a tight coupling between reasoning and trajectory generation. However, the robustness of such systems under realistic input perturbations remains largely unexplored. We show that these models are highly vulnerable to realistic input perturbations, achieving up to 89% attack success rate (ASR) on reasoning and up to 72% on trajectory manipulation in closed-loop simulation, leading to increased collision rates and degraded safety metrics. Using NVIDIA's recent Alpamayo models as representative industry-developed VLAs, we conduct the first systematic black-box study of reasoning-enabled VLA models under realistic textual input corruptions, evaluating their impact on reasoning and driving behavior. We introduce a reasoning-aware evaluation framework capturing both semantic and structural aspects of reasoning, along with safety-centric measures. We also introduce a benchmark for evaluating attacks and defenses on reasoning-trajectory interactions in autonomous driving. Our results highlight the need for rigorous evaluation and improved defenses to ensure the safety of reasoning-enabled VLA systems in autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。