将视觉感知与强化学习结合,提升自动驾驶决策可解释性与安全性。
Hierarchical End-to-End Autonomous Driving: Integrating BEV Perception with Deep Reinforcement Learning
- 用鸟瞰图融合多传感器数据,直接提取环境特征用于强化学习
- 碰撞率降低20%,优于当前最优方法
- 适合关注自动驾驶可解释性与端到端系统设计的研究者
端到端自动驾驶为传统模块化流程提供了简化方案,将感知、预测与规划整合于单一框架中。尽管深度强化学习(DRL)在此领域逐渐兴起,现有方法常忽略DRL特征提取与感知之间的关键关联。本文通过将DRL的特征提取网络直接映射至感知阶段,实现语义分割层面的清晰可解释性。基于鸟瞰图(BEV)表示,提出一种新型基于DRL的端到端驾驶框架,利用多传感器输入构建对环境的统一三维理解。该BEV系统将关键环境特征提取并转化为DRL的高层抽象状态,以支持更明智的控制决策。大量实验表明,该方法不仅显著提升可解释性,还在自动驾驶控制任务中大幅超越现有先进方法,碰撞率降低20%。
原文摘要 · Abstract (English)
End-to-end autonomous driving offers a streamlined alternative to the traditional modular pipeline, integrating perception, prediction, and planning within a single framework. While Deep Reinforcement Learning (DRL) has recently gained traction in this domain, existing approaches often overlook the critical connection between feature extraction of DRL and perception. In this paper, we bridge this gap by mapping the DRL feature extraction network directly to the perception phase, enabling clearer interpretation through semantic segmentation. By leveraging Bird's-Eye-View (BEV) representations, we propose a novel DRL-based end-to-end driving framework that utilizes multi-sensor inputs to construct a unified three-dimensional understanding of the environment. This BEV-based system extracts and translates critical environmental features into high-level abstract states for DRL, facilitating more informed control. Extensive experimental evaluations demonstrate that our approach not only enhances interpretability but also significantly outperforms state-of-the-art methods in autonomous driving control tasks, reducing the collision rate by 20%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。