arXiv:2506.05473cs.CV2025-06被引 2

用稀疏查询流式建模动态驾驶场景,速度提升5.9倍且精度领先。

S2GO: Streaming Sparse Gaussian Occupancy Prediction

  • 用可时序传播的3D查询替代密集表示,实现轻量化在线建模
  • 在nuScenes和KITTI上达1.5点IoU提升,推理快5.9倍
  • 适合实时自动驾驶场景的高效三维占用预测,尤其关注动态变化

尽管基于稀疏查询的表示在感知任务中展现出高效与优秀性能,当前最先进的3D占用预测方法仍依赖体素或稠密高斯表示。然而,稠密表示效率低,难以捕捉驾驶场景的时序动态。与以往工作不同,我们通过在线流式方式将场景压缩为一组紧凑的3D查询,并随时间传播。这些查询在每个时间步解码为语义高斯。我们引入去噪渲染目标来指导查询及其构成的高斯,以有效捕捉场景几何。得益于高效的查询表示,S2GO在nuScenes和KITTI占用基准上达到最优性能,相比先前方法(如GaussianWorld)提升1.5 IoU,推理速度提升5.9倍。

原文摘要 · Abstract (English)

Despite the demonstrated efficiency and performance of sparse query-based representations for perception, state-of-the-art 3D occupancy prediction methods still rely on voxel-based or dense Gaussian-based 3D representations. However, dense representations are slow, and they lack flexibility in capturing the temporal dynamics of driving scenes. Distinct from prior work, we instead summarize the scene into a compact set of 3D queries which are propagated through time in an online, streaming fashion. These queries are then decoded into semantic Gaussians at each timestep. We couple our framework with a denoising rendering objective to guide the queries and their constituent Gaussians in effectively capturing scene geometry. Owing to its efficient, query-based representation, S2GO achieves state-of-the-art performance on the nuScenes and KITTI occupancy benchmarks, outperforming prior art (e.g., GaussianWorld) by 1.5 IoU with 5.9x faster inference.

3D占用预测稀疏查询实时建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。