arXiv:2608.28656cs.ROcs.CV2026-08

提升自动驾驶在红绿灯路口的规则遵守能力,减少误刹误行。

RedLight-VLA: Models for traffic-rule grounding and behavioral emphasis in driving policies

  • 通过行为重加权和辅助视觉头,强化对红绿灯与停车线状态的感知。
  • 红灯停线越界率降至6.8%,速度误差降低12.7%,预测更精准。
  • 适合关注交通规则理解与驾驶行为鲁棒性的自动驾驶研究者。

行为克隆的视觉-语言-动作(VLA)驾驶策略在信号灯交叉口的罕见规则操作中表现不佳。制动与启动样本对平均轨迹损失贡献小,融合表征也缺乏对交通灯与停车线状态的显式监督。我们提出RedLight-VLA,采用专家未来轨迹和自动生成的感知目标,无需额外人工规则标注。首先,基于轨迹的行为重加权(BR)利用旋转不变的纵向动力学和保尺度压缩,强调罕见减速与加速,关闭时可完全还原基线。其次,并行辅助(AUX)头在连续融合后的规则令牌中显式建模交通灯与停车线状态,无需自回归语言生成或修改轨迹解码器。在20秒序列上评估,预测窗口5秒。控制变量保持相同骨干网络、训练数据、解码器和测试集。相比同构基线,RedLight-VLA将红灯停线越界率从7.3%降至6.8%,停线速度误差降低12.7%,3秒交通灯切片的ADE/FDE由0.274/0.964米降至0.247/0.897米。绿灯误停率从3.2%升至3.9%;但结合BR与AUX可缓解仅用AUX时更大的上升(4.0%)。联合模型还将非交通灯相关ADE/FDE从0.268/0.956米降至0.241/0.876米,在四个切片位移指标上均优于单一机制。

原文摘要 · Abstract (English)

Behavior-cloned Vision-Language-Action (VLA) driving policies struggle with rare rule-governed maneuvers at signalized intersections. Braking and launching examples contribute little to averaged trajectory loss, while fused representations lack explicit supervision for the governing traffic-light and stop-line state. We present RedLight-VLA, a training objective that uses expert futures and automatically generated perception targets without additional manual rule annotation. First, trajectory-derived behavioral reweighting (BR) emphasizes rare deceleration and acceleration using rotation-invariant longitudinal dynamics and a scale-preserving reduction that exactly recovers the baseline when disabled. Second, parallel auxiliary (AUX) heads ground traffic-light and stop-line state in continuous post-fusion rule tokens, without autoregressive language generation or changes to the trajectory decoder. We evaluate on a curated set of 20 s sequences with a 5 s prediction horizon. Controlled variants share the same backbone, training data, decoder, and evaluation population. Against an otherwise identical VLA baseline, RedLight-VLA reduces red-light stop-line overshoot from 7.3% to6.8%, reduces stop-line velocity error by 12.7%, and improves 3 s trafficlight-sliced ADE/FDE from 0.274/0.964 m to 0.247/0.897 m. Green-light false stops increase from 3.2% to 3.9%; however, combining BR with AUX supervision mitigates the larger increase observed for AUX alone (4.0%). The combined model also improves non-traffic-light ADE/FDE from 0.268/0.956 m to 0.241/0.876 m and outperforms either mechanism alone on all four sliced displacement measures.

自动驾驶视觉语言规则推理驾驶策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。