用双路径融合技术提升低密度4D雷达的自动驾驶感知能力
DRIFT: Dual-Representation Inter-Fusion Transformer for Automated Driving Perception with 4D Radar Point Clouds
- 设计双路径结构,分别提取局部细节与全局上下文特征
- 在VoD数据集上实现52.6% mAP,优于CenterPoint的45.4%
- 适合需要高鲁棒性、低成本感知方案的研发人员
4D雷达能提供带多普勒速度的三维点云,在恶劣天气下表现稳定且成本较低,但其点云密度远低于激光雷达。因此,需有效利用局部与全局场景信息。本文提出DRIFT模型,通过双路径架构融合局部与全局上下文:点路径捕获细粒度局部特征,柱状路径编码粗粒度全局特征。两条路径在多个阶段通过新型特征共享层交织,实现双重表示的充分融合。模型在广泛使用的View-of-Delft(VoD)数据集和内部私有数据集上评估,显著优于基线方法,在目标检测与自由道路估计任务中表现突出。例如,在VoD数据集上,DRIFT达到52.6%的均值平均精度(mAP),相较CenterPoint的45.4%有明显提升。
原文摘要 · Abstract (English)
4D radars, which provide 3D point cloud data along with Doppler velocity, are attractive components of modern automated driving systems due to their low cost and robustness under adverse weather conditions. However, they provide a significantly lower point cloud density than LiDAR sensors. This makes it important to exploit not only local but also global contextual scene information. This paper proposes DRIFT, a model that effectively captures and fuses both local and global contexts through a dual-path architecture. The model incorporates a point path to aggregate fine-grained local features and a pillar path to encode coarse-grained global features. These two parallel paths are intertwined via novel feature-sharing layers at multiple stages, enabling full utilization of both representations. DRIFT is evaluated on the widely used View-of-Delft (VoD) dataset and a proprietary internal dataset. It outperforms the baselines on the tasks of object detection and/or free road estimation. For example, DRIFT achieves a mean average precision (mAP) of 52.6% (compared to, say, 45.4% of CenterPoint) on the VoD dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。