arXiv:2504.12959cs.CV2025-04CVPR被引 12

提出GDFusion,用梯度下降统一融合多帧视觉信息,提升3D语义占据预测精度与效率。

Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy Prediction

论文配图:Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy Prediction
图 1 · 摘自论文原文
  • 将RNN重新解读为梯度下降,统一融合不同模态的时序特征
  • 在nuScenes上实现1.4%-4.8%的mIoU提升,内存降低27%-72%
  • 适用于自动驾驶中的3D语义占据预测,尤其关注时序一致性与运动校准

我们提出GDFusion,一种面向基于视觉的3D语义占据预测(VisionOcc)的时序融合方法。该方法深入探索VisionOcc框架中被忽视的时序融合机制,系统分析整个流程,识别出三个关键但未被重视的时序线索:场景级一致性、运动校准和几何补全。这些线索捕捉了时序演化的多样特性,并在VisionOcc各模块中发挥独特作用。为有效融合异构表示中的时序信号,我们提出一种新融合策略,通过重新诠释标准RNN的公式,利用特征上的梯度下降实现多元时序信息的统一整合,自然嵌入上述时序线索。在nuScenes数据集上的大量实验表明,GDFusion显著优于现有基线,在Occ3D基准上实现1.4%-4.8%的mIoU提升,同时内存消耗降低27%-72%。

原文摘要 · Abstract (English)

We present GDFusion, a temporal fusion method for vision-based 3D semantic occupancy prediction (VisionOcc). GDFusion opens up the underexplored aspects of temporal fusion within the VisionOcc framework, focusing on both temporal cues and fusion strategies. It systematically examines the entire VisionOcc pipeline, identifying three fundamental yet previously overlooked temporal cues: scene-level consistency, motion calibration, and geometric complementation. These cues capture diverse facets of temporal evolution and make distinct contributions across various modules in the VisionOcc framework. To effectively fuse temporal signals across heterogeneous representations, we propose a novel fusion strategy by reinterpreting the formulation of vanilla RNNs. This reinterpretation leverages gradient descent on features to unify the integration of diverse temporal information, seamlessly embedding the proposed temporal cues into the network. Extensive experiments on nuScenes demonstrate that GDFusion significantly outperforms established baselines. Notably, on Occ3D benchmark, it achieves 1.4\%-4.8\% mIoU improvements and reduces memory consumption by 27\%-72\%.

3D占据预测时序融合视觉感知自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。