融合视觉语言模型与端到端网络,提升自动驾驶轨迹预测精度
SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous Driving
- 共享视觉编码器实现视觉语言与端到端模型的特征级协同
- 采用渐进式轨迹推理机制,显著降低预测不确定性
- 在nuScenes数据集上实现更精准的轨迹预测,适合自动驾驶研发者
将视觉语言模型(VLM)融入自动驾驶系统,有助于缓解学习复杂性、可解释性及常识推理等关键挑战。然而,现有方法常因计算开销大,在高效集成与实时决策方面表现不佳。本文提出SOLVE框架,通过共享视觉编码器实现VLM与端到端(E2E)模型的特征级知识共享,促进二者深度协同。引入轨迹链式思维(T-CoT)机制,逐步优化轨迹预测,有效减少不确定性并提升准确性。结合时间解耦策略,使高质量的VLM输出与E2E模型的实时性能高效匹配。在nuScenes数据集上的评估表明,该方法在轨迹预测精度上取得显著提升,为构建更鲁棒可靠的自动驾驶系统提供了新路径。
原文摘要 · Abstract (English)
The integration of Vision-Language Models (VLMs) into autonomous driving systems has shown promise in addressing key challenges such as learning complexity, interpretability, and common-sense reasoning. However, existing approaches often struggle with efficient integration and realtime decision-making due to computational demands. In this paper, we introduce SOLVE, an innovative framework that synergizes VLMs with end-to-end (E2E) models to enhance autonomous vehicle planning. Our approach emphasizes knowledge sharing at the feature level through a shared visual encoder, enabling comprehensive interaction between VLM and E2E components. We propose a Trajectory Chain-of-Thought (T-CoT) paradigm, which progressively refines trajectory predictions, reducing uncertainty and improving accuracy. By employing a temporal decoupling strategy, SOLVE achieves efficient cooperation by aligning high-quality VLM outputs with E2E real-time performance. Evaluated on the nuScenes dataset, our method demonstrates significant improvements in trajectory prediction accuracy, paving the way for more robust and reliable autonomous driving systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。