融合多帧相机与4D雷达数据,提升复杂环境下的3D目标检测精度。
M^3Detection: Multi-Frame Multi-Level Feature Fusion for Multi-Modal 3D Object Detection with Camera and 4D Imaging Radar
- 通过多帧多层级特征融合,整合相机语义与雷达时空信息。
- 在VoD和TJ4DRadSet上达到当前最优性能,召回率提升6.2%。
- 适合自动驾驶感知系统,尤其适用于恶劣天气场景。
4D成像雷达在恶劣天气下具备鲁棒感知能力,而相机传感器提供丰富的语义信息。二者融合具有实现低成本3D感知的巨大潜力。然而,现有相机-雷达融合方法大多仅限于单帧输入,仅捕捉场景的部分视图。不完整的场景信息,叠加图像退化和4D雷达点云稀疏性,制约了整体检测性能。相比之下,多帧融合能提供更丰富的时空信息,但面临两大挑战:如何在帧与模态间实现稳健有效的特征融合,以及如何降低冗余特征提取带来的计算开销。为此,我们提出M^3Detection,一个统一的多帧3D目标检测框架,对来自相机和4D成像雷达的多模态数据进行多层次特征融合。该框架利用基线检测器的中间特征,并结合追踪器生成参考轨迹,提升计算效率并为第二阶段提供更丰富信息。在第二阶段,设计了基于雷达信息引导的全局级跨目标特征聚合模块,用于对齐候选框间的全局特征;同时引入局部级跨网格特征聚合模块,沿参考轨迹扩展局部特征,增强细粒度目标表征。聚合后的特征由轨迹级多帧时空推理模块处理,以编码跨帧交互并强化时序表示。在VoD和TJ4DRadSet数据集上的大量实验表明,M^3Detection实现了当前最优的3D检测性能,验证了其在多帧相机-4D成像雷达融合检测中的有效性。
原文摘要 · Abstract (English)
Recent advances in 4D imaging radar have enabled robust perception in adverse weather, while camera sensors provide dense semantic information. Fusing the these complementary modalities has great potential for cost-effective 3D perception. However, most existing camera-radar fusion methods are limited to single-frame inputs, capturing only a partial view of the scene. The incomplete scene information, compounded by image degradation and 4D radar sparsity, hinders overall detection performance. In contrast, multi-frame fusion offers richer spatiotemporal information but faces two challenges: achieving robust and effective object feature fusion across frames and modalities, and mitigating the computational cost of redundant feature extraction. Consequently, we propose M^3Detection, a unified multi-frame 3D object detection framework that performs multi-level feature fusion on multi-modal data from camera and 4D imaging radar. Our framework leverages intermediate features from the baseline detector and employs the tracker to produce reference trajectories, improving computational efficiency and providing richer information for second-stage. In the second stage, we design a global-level inter-object feature aggregation module guided by radar information to align global features across candidate proposals and a local-level inter-grid feature aggregation module that expands local features along the reference trajectories to enhance fine-grained object representation. The aggregated features are then processed by a trajectory-level multi-frame spatiotemporal reasoning module to encode cross-frame interactions and enhance temporal representation. Extensive experiments on the VoD and TJ4DRadSet datasets demonstrate that M^3Detection achieves state-of-the-art 3D detection performance, validating its effectiveness in multi-frame detection with camera-4D imaging radar fusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。