用视觉语言模型指导自动驾驶,提升长尾场景应对能力。
SteerVLA: Steering Vision-Language-Action Models in Long-Tail Driving Scenarios
- 用VLM生成细粒度语言指令,引导低层驾驶策略
- 在长尾场景中驾驶得分提升8.04分,整体提升4.77分
- 适合研究自动驾驶决策与多模态控制融合的学者
自动驾驶的核心挑战在于将高层语义推理与底层反应式控制相结合。虽然大规模视觉语言模型(VLM)具备强大的常识推理能力,但缺乏安全车辆控制所需的具身经验。本文提出SteerVLA,利用VLM的推理能力生成细粒度语言指令,以引导视觉-语言-动作(VLA)驾驶策略。关键在于高阶VLM与低阶VLA之间丰富的语言接口,使高层推理能更好地对接控制输出。为提供与车辆控制对齐的细粒度语言监督,我们使用VLM为现有驾驶数据添加详细语言标注,实证其对有效推理和可操控性至关重要。在封闭环路基准测试中,SteerVLA在整体驾驶评分上优于现有方法4.77分,在长尾子集上提升8.04分。
原文摘要 · Abstract (English)
A fundamental challenge in autonomous driving is the integration of high-level, semantic reasoning for long-tail events with low-level, reactive control for robust driving. While large vision-language models (VLMs) trained on web-scale data offer powerful common-sense reasoning, they lack the grounded experience necessary for safe vehicle control. We posit that an effective autonomous agent should leverage the world knowledge of VLMs to guide a steerable driving policy toward robust control in driving scenarios. To this end, we propose SteerVLA, which leverages the reasoning capabilities of VLMs to produce fine-grained language instructions that steer a vision-language-action (VLA) driving policy. Key to our method is this rich language interface between the high-level VLM and low-level VLA, which allows the high-level policy to more effectively ground its reasoning in the control outputs of the low-level policy. To provide fine-grained language supervision aligned with vehicle control, we leverage a VLM to augment existing driving data with detailed language annotations, which we find to be essential for effective reasoning and steerability. We evaluate SteerVLA on a challenging closed-loop benchmark, where it outperforms state-of-the-art methods by 4.77 points in overall driving score and by 8.04 points on a long-tail subset. The project website is available at: https://steervla.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。