arXiv:2604.05378cs.CLcs.CV2026-04被引 2

测试语言指令微小变化对自动驾驶模型的影响,发现微调指令会引发严重性能下降。

ICR-Drive: Instruction Counterfactual Robustness for End-to-End Language-Driven Autonomous Driving

  • 构建四类指令扰动:改写、模糊、噪声、误导性指令
  • 在CARLA仿真中重播相同路线,量化指令变化导致的性能退化
  • 揭示当前自动驾驶模型在真实指令场景下的可靠性短板

视觉-语言-动作(VLA)模型的进步使得语言驱动的自动驾驶系统能在闭环仿真中执行自然语言导航指令,但现有评估大多假设指令精确且结构完整。实际部署中,指令常存在表述差异、信息缺失,甚至包含误导性或权威性措辞,导致指令鲁棒性未被充分衡量。本文提出ICR-Drive,一个用于端到端语言驱动自动驾驶中指令反事实鲁棒性的诊断框架。该框架生成涵盖四类扰动的指令变体:改写、模糊、噪声和误导性,其中误导性指令与导航目标冲突并试图覆盖原意图。在相同CARLA路线、模拟器配置和种子下重播,以隔离性能变化仅由指令语言引起。使用标准CARLA Leaderboard指标及各族扰动相对于基线指令的性能下降来量化鲁棒性。在LMDrive和BEVDriver上的实验表明,轻微指令变化可导致显著性能下降并引发不同故障模式,暴露出将具身基础模型应用于高安全性驾驶任务时的可靠性差距。

原文摘要 · Abstract (English)

Recent progress in vision-language-action (VLA) models has enabled language-conditioned driving agents to execute natural-language navigation commands in closed-loop simulation, yet standard evaluations largely assume instructions are precise and well-formed. In deployment, instructions vary in phrasing and specificity, may omit critical qualifiers, and can occasionally include misleading, authority-framed text, leaving instruction-level robustness under-measured. We introduce ICR-Drive, a diagnostic framework for instruction counterfactual robustness in end-to-end language-conditioned autonomous driving. ICR-Drive generates controlled instruction variants spanning four perturbation families: Paraphrase, Ambiguity, Noise, and Misleading, where Misleading variants conflict with the navigation goal and attempt to override intent. We replay identical CARLA routes under matched simulator configurations and seeds to isolate performance changes attributable to instruction language. Robustness is quantified using standard CARLA Leaderboard metrics and per-family performance degradation relative to the baseline instruction. Experiments on LMDrive and BEVDriver show that minor instruction changes can induce substantial performance drops and distinct failure modes, revealing a reliability gap for deploying embodied foundation models in safety-critical driving.

自动驾驶语言模型鲁棒性评测CARLA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。