让自动驾驶像人一样分层思考,提升决策准确性和安全性。
ReAL-AD: Towards Human-Like Reasoning in End-to-End Autonomous Driving
- 基于人类驾驶三层次认知模型,用视觉语言模型增强环境理解
- 引入分层推理模块,使规划结果比现有方法准确率提升30%以上
- 适合关注可解释性与类人决策的自动驾驶研究者和工程师
端到端自动驾驶虽能统一感知、预测与规划,减少信息损失并增强适应性,但现有方法多依赖固定稀疏轨迹监督,难以捕捉人类驾驶员自然具备的分层推理过程。为此,我们提出ReAL-AD——一种基于人类认知三层模型(驾驶策略、驾驶决策、驾驶操作)的推理增强学习框架。通过引入视觉语言模型(VLMs)提升情境感知与结构化推理能力,设计了三个核心组件:战略推理注入器,基于VLM生成的上下文洞察制定高层驾驶策略;战术推理整合器,将战略意图转化为可解释的战术选择(如变道、超车、调速);分层轨迹解码器,逐步将战术决策转为精准控制动作,实现平滑且类人的轨迹执行。大量实验表明,该框架使规划准确率与安全性提升超过30%,显著增强了端到端自动驾驶的可解释性与类人推理能力。
原文摘要 · Abstract (English)
End-to-end autonomous driving has emerged as a promising approach to unify perception, prediction, and planning within a single framework, reducing information loss and improving adaptability. However, existing methods often rely on fixed and sparse trajectory supervision, limiting their ability to capture the hierarchical reasoning process that human drivers naturally employ. To bridge this gap, we propose ReAL-AD, a Reasoning-Augmented Learning framework that structures decision-making in autonomous driving based on the three-tier human cognitive model: Driving Strategy, Driving Decision, and Driving Operation, where Vision-Language Models (VLMs) are incorporated to enhance situational awareness and structured reasoning across these levels. Specifically, we introduce: (1) the Strategic Reasoning Injector, which formulates high-level driving strategies by interpreting complex traffic contexts from VLM-generated insights; (2) the Tactical Reasoning Integrator, which refines strategic intent into interpretable tactical choices such as lane changes, overtaking, and speed adjustments; and (3) the Hierarchical Trajectory Decoder, which progressively translates tactical decisions into precise control actions for smooth and human-like trajectory execution. Extensive evaluations show that integrating our framework improves planning accuracy and safety by over 30%, making end-to-end autonomous driving more interpretable and aligned with human-like hierarchical reasoning. The project page can be found at: \href{https://4dvlab.github.io/project_page/realad}{\texttt{4dvlab.github.io/project\_page/realad}}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。