用视觉语言模型实现可解释的闭环自动驾驶,性能超越现有方法。
X-Driver: Explainable Autonomous Driving with Vision-Language Models
- 引入多模态大模型结合思维链与自回归推理,统一感知与决策
- 在CARLA仿真中多个基准上表现优于当前最先进水平,闭环成功率更高
- 提升驾驶决策可解释性,适合关注透明化自动驾驶的研究者
端到端自动驾驶已取得显著进展,具有系统简化和开环/闭环性能更优的优势。然而,现有框架在闭环评估中仍存在成功率低的问题,制约其真实部署。本文提出X-Driver,一个基于多模态大模型(MLLMs)的统一闭环自动驾驶框架,采用思维链(CoT)与自回归建模增强感知与决策能力。我们在CARLA仿真环境中的多个公开基准(包括Bench2Drive[6])上验证了X-Driver,实验结果表明其在闭环任务中表现优于当前最先进方法,同时提升了驾驶决策的可解释性。研究强调了结构化推理在端到端自动驾驶中的重要性,并为未来闭环自动驾驶研究提供了强有力的基线。
原文摘要 · Abstract (English)
End-to-end autonomous driving has advanced significantly, offering benefits such as system simplicity and stronger driving performance in both open-loop and closed-loop settings than conventional pipelines. However, existing frameworks still suffer from low success rates in closed-loop evaluations, highlighting their limitations in real-world deployment. In this paper, we introduce X-Driver, a unified multi-modal large language models(MLLMs) framework designed for closed-loop autonomous driving, leveraging Chain-of-Thought(CoT) and autoregressive modeling to enhance perception and decision-making. We validate X-Driver across multiple autonomous driving tasks using public benchmarks in CARLA simulation environment, including Bench2Drive[6]. Our experimental results demonstrate superior closed-loop performance, surpassing the current state-of-the-art(SOTA) while improving the interpretability of driving decisions. These findings underscore the importance of structured reasoning in end-to-end driving and establish X-Driver as a strong baseline for future research in closed-loop autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。