解决自动驾驶中3D检测的时空不一致问题,提升感知稳定性
Co-Fusion4D: Spatio-temporal Collaborative Fusion for Robust 3D Object Detection

- 以当前帧为中心,融合历史帧前做时空对齐与筛选
- 在nuScenes上达74.9% mAP、75.6% NDS,无需测试增强
- 双注意力融合机制有效抑制噪声传播,增强运动一致性
在自动驾驶中,3D目标检测对精准感知和可靠决策至关重要。然而,目标运动与自车运动常导致基于鸟瞰图(BEV)检测器出现跨帧时空不一致,引发时间序列特征错位,降低时空一致性。为此,本文提出Co-Fusion4D统一框架,显式保持跨帧时空一致性并抑制时序特征漂移。该框架采用当前帧为中心策略,将当前帧作为主要信息源,经时空过滤与对齐后选择性融合历史帧。这种主导-互补机制有效缓解累积对齐误差,抑制噪声传播,并利用可靠时序线索构建更一致的BEV表示。此外,引入双注意力融合(DAF)模块,联合使用帧内空间注意力与帧间时间注意力,自适应对齐与融合多帧特征,突出运动一致区域,抑制虚假关联。相比传统均匀融合范式,该设计显著提升BEV表示的时间稳定性与判别能力。在nuScenes基准上大量实验表明,Co-Fusion4D达到74.9% mAP与75.6% NDS,且无需测试时增强或外部数据。
原文摘要 · Abstract (English)
In autonomous driving, 3D object detection is essential for accurate perception and reliable decision-making. However, object motion and ego-motion often induce cross-frame spatiotemporal inconsistencies in BEV-based detectors, leading to temporal BEV feature misalignment and degraded spatiotemporal consistency. To address these challenges, we propose Co-Fusion4D, a unified framework that explicitly preserves cross-frame spatiotemporal consistency and suppresses temporal feature drift. Co-Fusion4D adopts a current-frame-centric strategy, treating the current frame as the primary source of information while selectively incorporating historical frames after spatiotemporal filtering and alignment. This dominant-complementary mechanism effectively mitigates cumulative alignment errors, suppresses noisy feature propagation, and exploits reliable temporal cues for a more consistent BEV representation. In addition, Co-Fusion4D integrates a Dual Attention Fusion (DAF) module to further enhance spatiotemporal feature interaction. DAF jointly leverages intra-frame spatial attention and inter-frame temporal attention to adaptively align and fuse multi-frame features, emphasizing motion-consistent regions while suppressing spurious correlations. By departing from conventional uniform fusion paradigms, this design substantially improves the temporal stability and discriminative capability of BEV representations. Extensive experiments on the nuScenes benchmark demonstrate that Co-Fusion4D achieves state-of-the-art performance, with 74.9% mAP and 75.6% NDS, without relying on test-time augmentation or external data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。