用视觉语言模型提升自动驾驶轨迹预测的决策能力
SimpleVSF: VLM-Scoring Fusion for Trajectory Prediction of End-to-End Autonomous Driving
- 结合传统评分与视觉语言模型增强评分
- 通过量化融合与上下文感知融合提升规划效果
- 适合关注智能驾驶决策优化的研究者
端到端自动驾驶已成为实现鲁棒智能驾驶策略的有前景范式,但现有方法在复杂场景中仍面临决策不优的问题。本文提出SimpleVSF(Simple VLM-Scoring Fusion)框架,通过利用视觉语言模型(VLMs)的认知能力与先进轨迹融合技术,增强端到端规划。采用传统评分器与新型VLM增强评分器,并引入稳健的权重融合器进行量化聚合,以及强大的VLM-based融合器实现定性、上下文感知的决策。作为ICCV 2025 NAVSIM v2端到端驾驶挑战赛的领先方法,SimpleVSF在安全性、舒适性与效率之间实现了卓越平衡。
原文摘要 · Abstract (English)
End-to-end autonomous driving has emerged as a promising paradigm for achieving robust and intelligent driving policies. However, existing end-to-end methods still face significant challenges, such as suboptimal decision-making in complex scenarios. In this paper,we propose SimpleVSF (Simple VLM-Scoring Fusion), a novel framework that enhances end-to-end planning by leveraging the cognitive capabilities of Vision-Language Models (VLMs) and advanced trajectory fusion techniques. We utilize the conventional scorers and the novel VLM-enhanced scorers. And we leverage a robust weight fusioner for quantitative aggregation and a powerful VLM-based fusioner for qualitative, context-aware decision-making. As the leading approach in the ICCV 2025 NAVSIM v2 End-to-End Driving Challenge, our SimpleVSF framework demonstrates state-of-the-art performance, achieving a superior balance between safety, comfort, and efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。