arXiv:2507.20963cs.CV2025-07

通过全局时序聚合提升3D占位预测的环境感知能力

GTAD: Global Temporal Aggregation Denoising Learning for 3D Semantic Occupancy Prediction

  • 引入全局时序聚合框架,融合当前与历史帧的时序特征
  • 在nuScenes和Occ3D-nuScenes上实现更优的占位预测性能
  • 适合需要精准动态环境理解的自动驾驶系统

准确感知动态环境是自动驾驶与机器人系统的基本任务。现有方法主要依赖相邻帧间的局部时序交互,未能有效利用全局序列信息。为此,我们提出一种名为GTAD的全局时序聚合去噪网络,构建了一种全新的整体3D场景理解范式。该方法在模型内引入潜在去噪网络,聚合当前时刻的局部时序特征与历史序列的全局时序特征,从而有效感知邻近帧的细粒度时序信息及历史观测中的全局时序模式,实现对环境更连贯、全面的理解。在nuScenes与Occ3D-nuScenes基准上的大量实验与消融研究验证了方法的优势。

原文摘要 · Abstract (English)

Accurately perceiving dynamic environments is a fundamental task for autonomous driving and robotic systems. Existing methods inadequately utilize temporal information, relying mainly on local temporal interactions between adjacent frames and failing to leverage global sequence information effectively. To address this limitation, we investigate how to effectively aggregate global temporal features from temporal sequences, aiming to achieve occupancy representations that efficiently utilize global temporal information from historical observations. For this purpose, we propose a global temporal aggregation denoising network named GTAD, introducing a global temporal information aggregation framework as a new paradigm for holistic 3D scene understanding. Our method employs an in-model latent denoising network to aggregate local temporal features from the current moment and global temporal features from historical sequences. This approach enables the effective perception of both fine-grained temporal information from adjacent frames and global temporal patterns from historical observations. As a result, it provides a more coherent and comprehensive understanding of the environment. Extensive experiments on the nuScenes and Occ3D-nuScenes benchmark and ablation studies demonstrate the superiority of our method.

3D占位预测时序建模自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。