arXiv:2504.19749cs.CV2025-04CVPR被引 25

通过显式状态修复提升3D占据与场景流预测精度,降低内存占用。

STCOcc: Sparse Spatial-Temporal Cascade Renovation for 3D Occupancy and Scene Flow Prediction

论文配图:STCOcc: Sparse Spatial-Temporal Cascade Renovation for 3D Occupancy and Scene Flow Prediction
图 1 · 摘自论文原文
  • 基于占据状态的稀疏注意力机制,分层重构3D特征。
  • 在KITTI和Waymo数据集上实现更高精度的占据与场景流预测。
  • 适合需要高效高精度3D动态场景建模的研究者使用。

3D占据与场景流提供了对三维场景的详细动态表征。针对3D空间的稀疏性与复杂性,现有视觉中心方法采用基于隐式学习的方法建模时空信息,但难以捕捉局部细节,削弱了空间区分能力。为此,我们提出一种新型显式状态建模方法,利用占据状态对3D特征进行重构。具体而言,设计了一种稀疏遮挡感知注意力机制,并结合级联优化策略,借助占据状态信息精确重构3D特征。此外,引入一种新方法建模长期动态交互,在降低计算成本的同时保持空间信息。相比先前最先进方法,本方法在占据和场景流预测上的RayIoU与mAVE指标均更优,且训练时GPU内存降至8.7GB。代码已开源。

原文摘要 · Abstract (English)

3D occupancy and scene flow offer a detailed and dynamic representation of 3D scene. Recognizing the sparsity and complexity of 3D space, previous vision-centric methods have employed implicit learning-based approaches to model spatial and temporal information. However, these approaches struggle to capture local details and diminish the model's spatial discriminative ability. To address these challenges, we propose a novel explicit state-based modeling method designed to leverage the occupied state to renovate the 3D features. Specifically, we propose a sparse occlusion-aware attention mechanism, integrated with a cascade refinement strategy, which accurately renovates 3D features with the guidance of occupied state information. Additionally, we introduce a novel method for modeling long-term dynamic interactions, which reduces computational costs and preserves spatial information. Compared to the previous state-of-the-art methods, our efficient explicit renovation strategy not only delivers superior performance in terms of RayIoU and mAVE for occupancy and scene flow prediction but also markedly reduces GPU memory usage during training, bringing it down to 8.7GB. Our code is available on https://github.com/lzzzzzm/STCOcc

3D占据场景流显式建模稀疏注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。