arXiv:2506.09417cs.CV2025-06NeurIPS被引 7

用双高斯稀疏表示,提升自动驾驶场景占用预测精度与效率

ODG: Occupancy Prediction Using Dual Gaussians

论文配图:ODG: Occupancy Prediction Using Dual Gaussians
图 1 · 摘自论文原文
  • 分动静两部分建模,用双高斯稀疏查询捕捉复杂场景动态
  • 在Occ3D-nuScenes和Waymo上达到新SOTA,推理开销低
  • 结合渲染监督,实现像素级对齐,适合自动驾驶感知任务

占用预测从周围环境的摄像头图像中推断精细的三维几何与语义信息,是自动驾驶的关键感知任务。现有方法要么采用密集网格表示,难以扩展到高分辨率;要么仅用一组稀疏查询学习整个场景,无法充分处理物体多样性。本文提出ODG,一种分层双稀疏高斯表示,有效捕捉复杂场景动态。基于驾驶场景可普遍分解为静态与动态部分的观察,定义双高斯查询以更好建模多样物体。采用分层高斯变换器预测占据体素中心、语义类别及高斯参数。利用3D高斯溅射的实时渲染能力,引入深度与语义图标注的渲染监督,实现像素级对齐以增强占用学习。在Occ3D-nuScenes和Occ3D-Waymo基准上的大量实验表明,该方法达到新的最优性能,同时保持低推理成本。

原文摘要 · Abstract (English)

Occupancy prediction infers fine-grained 3D geometry and semantics from camera images of the surrounding environment, making it a critical perception task for autonomous driving. Existing methods either adopt dense grids as scene representation, which is difficult to scale to high resolution, or learn the entire scene using a single set of sparse queries, which is insufficient to handle the various object characteristics. In this paper, we present ODG, a hierarchical dual sparse Gaussian representation to effectively capture complex scene dynamics. Building upon the observation that driving scenes can be universally decomposed into static and dynamic counterparts, we define dual Gaussian queries to better model the diverse scene objects. We utilize a hierarchical Gaussian transformer to predict the occupied voxel centers and semantic classes along with the Gaussian parameters. Leveraging the real-time rendering capability of 3D Gaussian Splatting, we also impose rendering supervision with available depth and semantic map annotations injecting pixel-level alignment to boost occupancy learning. Extensive experiments on the Occ3D-nuScenes and Occ3D-Waymo benchmarks demonstrate our proposed method sets new state-of-the-art results while maintaining low inference cost.

占用预测高斯表示自动驾驶3D感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。