用视觉语言模型提升自动驾驶决策能力,实现端到端闭环表现突破。
ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation

- 融合视觉语言模型与长期历史感知,生成精准驾驶轨迹
- 在Bench2Drive上达77.74分驾驶得分,成功率54.62%,领先超19%和14分
- 首次统一语义推理与动作输出空间,适合复杂交互场景研究者
端到端自动驾驶方法在交互式闭环评估中仍因因果推理能力有限而表现不佳。现有方法尝试利用视觉语言模型(VLM)的强理解与推理能力解决此问题,但多数VLM在闭环评估中表现不佳,根源在于语义推理空间与纯数值轨迹输出空间之间的鸿沟。为此,我们提出ORION——一种通过视觉语言指令生成动作的完整端到端自动驾驶框架。ORION创新性地结合了QT-Former以聚合长期历史上下文、大语言模型(LLM)进行驾驶场景推理,以及生成式规划器实现高精度轨迹预测。同时,通过对齐推理空间与动作空间,实现视觉问答(VQA)与规划任务的统一端到端优化。该方法在Bench2Drive挑战数据集上取得77.74分驾驶得分(DS)和54.62%成功率(SR),显著超越当前最优方法,分别提升14.28分与19.61个百分点。
原文摘要 · Abstract (English)
End-to-end (E2E) autonomous driving methods still struggle to make correct decisions in interactive closed-loop evaluation due to limited causal reasoning capability. Current methods attempt to leverage the powerful understanding and reasoning abilities of Vision-Language Models (VLMs) to resolve this dilemma. However, the problem is still open that few VLMs for E2E methods perform well in the closed-loop evaluation due to the gap between the semantic reasoning space and the purely numerical trajectory output in the action space. To tackle this issue, we propose ORION, a holistic E2E autonomous driving framework by vision-language instructed action generation. ORION uniquely combines a QT-Former to aggregate long-term history context, a Large Language Model (LLM) for driving scenario reasoning, and a generative planner for precision trajectory prediction. ORION further aligns the reasoning space and the action space to implement a unified E2E optimization for both visual question-answering (VQA) and planning tasks. Our method achieves an impressive closed-loop performance of 77.74 Driving Score (DS) and 54.62% Success Rate (SR) on the challenge Bench2Drive datasets, which outperforms state-of-the-art (SOTA) methods by a large margin of 14.28 DS and 19.61% SR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。