arXiv:2604.01723cs.ROcs.AI2026-04被引 4

让自动驾驶模型理解指令与环境的因果关系,提升安全决策能力。

Causal Scene Narration with Runtime Safety Supervision for Vision-Language-Action Driving

  • 用因果结构重排文本输入,让指令与环境约束对齐
  • 闭环测试中驾驶得分提升超31%,且抗感知噪声能力强
  • 适合研究视觉语言动作系统安全性的研究人员

面向自动驾驶的视觉-语言-行动(VLA)模型需整合导航指令、危险警告和交通状态描述等多源文本,但现有系统常将这些信息割裂处理,迫使模型自行判断哪些环境约束与当前操作相关。本文提出因果场景叙述(CSN),在推理时通过意图-约束对齐、量化定位和结构化分离,零成本重构文本输入。结合基于Simplex的运行时安全监督及训练时使用Plackett-Luce DPO与负对数似然正则化的对齐策略。在多城镇闭环CARLA评估中,CSN使原始LMDrive的驾驶得分提升+31.1%,偏好对齐版本提升+24.5%。受控消融显示,因果结构贡献了39.1%的性能增益,其余来自信息量本身。感知噪声消融表明,该方法对真实传感误差具有鲁棒性。语义安全监督改善违规分数,而反应式碰撞时间监控则导致性能下降,说明VLA系统需要以意图为导向的监控机制。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models for autonomous driving must integrate diverse textual inputs, including navigation commands, hazard warnings, and traffic state descriptions, yet current systems often present these as disconnected fragments, forcing the model to discover on its own which environmental constraints are relevant to the current maneuver. We introduce Causal Scene Narration (CSN), which restructures VLA text inputs through intent-constraint alignment, quantitative grounding, and structured separation, at inference time with zero GPU cost. We complement CSN with Simplex-based runtime safety supervision and training-time alignment via Plackett-Luce DPO with negative log-likelihood (NLL) regularization. A multi-town closed-loop CARLA evaluation shows that CSN improves Driving Score by +31.1% on original LMDrive and +24.5% on the preference-aligned variant. A controlled ablation reveals that causal structure accounts for 39.1% of this gain, with the remainder attributable to information content alone. A perception noise ablation confirms that CSN's benefit is robust to realistic sensing errors. Semantic safety supervision improves Infraction Score, while reactive Time-To-Collision monitoring degrades performance, demonstrating that intent-aware monitoring is needed for VLA systems.

自动驾驶因果推理VLA模型安全监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。