融合多视角4D雷达与摄像头,实现全天候精准3D目标检测与占位预测。
Doracamom: Joint 3D Detection and Occupancy Prediction with Multi-view 4D Radars and Cameras for Omnidirectional Perception
- 用雷达几何先验和图像语义特征初始化体素查询,提升感知基础
- 通过双分支时序编码器在BEV和体素空间并行学习时空特征
- 跨模态注意力融合互补信息,辅助任务增强特征质量
3D目标检测与占位预测是自动驾驶中的关键任务,近年来视觉方法虽有进展,但在恶劣环境下仍受限。将摄像头与新一代4D成像雷达融合以实现统一多任务感知具有重要意义,但相关研究仍较少。本文提出Doracamom,首个融合多视角相机与4D雷达的联合3D目标检测与语义占位预测框架,实现全方位环境感知。我们设计了新型粗粒度体素查询生成器,结合4D雷达几何先验与图像语义特征初始化体素查询,为后续Transformer精修奠定坚实基础。为利用时间信息,提出双分支时序编码器,在俯视图(BEV)与体素空间并行处理多模态时序特征,实现全面时空表征学习。进一步提出跨模态BEV-体素融合模块,通过注意力机制自适应融合互补特征,并引入辅助任务提升特征质量。在OmniHD-Scenes、View-of-Delft(VoD)和TJ4DRadSet数据集上的大量实验表明,Doracamom在两项任务上均达到领先性能,树立了多模态3D感知新基准。代码与模型将公开。
原文摘要 · Abstract (English)
3D object detection and occupancy prediction are critical tasks in autonomous driving, attracting significant attention. Despite the potential of recent vision-based methods, they encounter challenges under adverse conditions. Thus, integrating cameras with next-generation 4D imaging radar to achieve unified multi-task perception is highly significant, though research in this domain remains limited. In this paper, we propose Doracamom, the first framework that fuses multi-view cameras and 4D radar for joint 3D object detection and semantic occupancy prediction, enabling comprehensive environmental perception. Specifically, we introduce a novel Coarse Voxel Queries Generator that integrates geometric priors from 4D radar with semantic features from images to initialize voxel queries, establishing a robust foundation for subsequent Transformer-based refinement. To leverage temporal information, we design a Dual-Branch Temporal Encoder that processes multi-modal temporal features in parallel across BEV and voxel spaces, enabling comprehensive spatio-temporal representation learning. Furthermore, we propose a Cross-Modal BEV-Voxel Fusion module that adaptively fuses complementary features through attention mechanisms while employing auxiliary tasks to enhance feature quality. Extensive experiments on the OmniHD-Scenes, View-of-Delft (VoD), and TJ4DRadSet datasets demonstrate that Doracamom achieves state-of-the-art performance in both tasks, establishing new benchmarks for multi-modal 3D perception. Code and models will be publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。