arXiv:2503.08612cs.ROcs.CV2025-03ICCV被引 38

提出统一框架,让自动驾驶模型闭环控制更精准。

HiP-AD: Hierarchical and Multi-Granularity Planning with Deformable Attention for Autonomous Driving in a Single Decoder

  • 用多粒度规划查询融合空间、时间与驾驶风格信息。
  • 通过可变形注意力从图像中动态提取物理位置特征。
  • 支持感知与规划在鸟瞰图中迭代交互,适合真实场景应用。

尽管端到端自动驾驶技术近年取得显著进展,但在闭环评估中仍表现不佳,规划在查询设计与交互中的潜力尚未充分挖掘。本文提出一种多粒度规划查询表示,融合异构路点(包括空间、时间及驾驶风格路点)并基于多种采样模式,为轨迹预测提供额外监督,提升自车闭环控制精度。同时,显式利用规划轨迹的几何特性,通过可变形注意力机制根据物理位置高效检索相关图像特征。结合上述策略,提出新型端到端自动驾驶框架HiP-AD,其在统一解码器中同步完成感知、预测与规划。该框架使规划查询可在鸟瞰图空间与感知查询迭代交互,并动态从透视视角提取图像特征。实验表明,HiP-AD在闭环基准Bench2Drive上超越所有现有端到端方法,在真实数据集nuScenes上也达到竞争力表现。

原文摘要 · Abstract (English)

Although end-to-end autonomous driving (E2E-AD) technologies have made significant progress in recent years, there remains an unsatisfactory performance on closed-loop evaluation. The potential of leveraging planning in query design and interaction has not yet been fully explored. In this paper, we introduce a multi-granularity planning query representation that integrates heterogeneous waypoints, including spatial, temporal, and driving-style waypoints across various sampling patterns. It provides additional supervision for trajectory prediction, enhancing precise closed-loop control for the ego vehicle. Additionally, we explicitly utilize the geometric properties of planning trajectories to effectively retrieve relevant image features based on physical locations using deformable attention. By combining these strategies, we propose a novel end-to-end autonomous driving framework, termed HiP-AD, which simultaneously performs perception, prediction, and planning within a unified decoder. HiP-AD enables comprehensive interaction by allowing planning queries to iteratively interact with perception queries in the BEV space while dynamically extracting image features from perspective views. Experiments demonstrate that HiP-AD outperforms all existing end-to-end autonomous driving methods on the closed-loop benchmark Bench2Drive and achieves competitive performance on the real-world dataset nuScenes.

自动驾驶端到端多模态规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。