arXiv:2509.16552cs.CVcs.RO2025-09中稿 · ICRA被引 4

用时空高斯点云提升视觉自动驾驶中的3D语义占位预测精度

ST-GS: Vision-Based 3D Semantic Occupancy Prediction with Spatial-Temporal Gaussian Splatting

  • 设计双模注意力机制增强高斯表示的空间交互
  • 在nuScenes上达到当前最优,且时间一致性显著提升
  • 适合关注多帧场景理解与高效建模的自动驾驶研究者

3D占位预测对视觉主导的自动驾驶中全面场景理解至关重要。近期工作虽采用3D语义高斯模型降低计算开销,但仍受限于多视角空间交互不足和多帧时间一致性差。为此,本文提出一种新型时空高斯点云(ST-GS)框架,以增强现有高斯基流水线中的时空建模能力。具体而言,我们设计了一种基于引导信息的空间聚合策略,集成于双模注意力机制中,强化高斯表示的空间交互。同时,引入几何感知的时间融合方案,有效利用历史上下文提升场景补全的时间连续性。在大规模nuScenes占位预测基准上的大量实验表明,所提方法不仅达到当前最优性能,相比现有高斯基方法在时间一致性上也有显著改善。

原文摘要 · Abstract (English)

3D occupancy prediction is critical for comprehensive scene understanding in vision-centric autonomous driving. Recent advances have explored utilizing 3D semantic Gaussians to model occupancy while reducing computational overhead, but they remain constrained by insufficient multi-view spatial interaction and limited multi-frame temporal consistency. To overcome these issues, in this paper, we propose a novel Spatial-Temporal Gaussian Splatting (ST-GS) framework to enhance both spatial and temporal modeling in existing Gaussian-based pipelines. Specifically, we develop a guidance-informed spatial aggregation strategy within a dual-mode attention mechanism to strengthen spatial interaction in Gaussian representations. Furthermore, we introduce a geometry-aware temporal fusion scheme that effectively leverages historical context to improve temporal continuity in scene completion. Extensive experiments on the large-scale nuScenes occupancy prediction benchmark showcase that our proposed approach not only achieves state-of-the-art performance but also delivers markedly better temporal consistency compared to existing Gaussian-based methods.

3D占位高斯点云自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。