arXiv:2503.18073cs.CVcs.RO2025-03被引 4

端到端实现开放词汇全景重建,提升3D场景理解精度与效率

PanopticSplatting: End-to-End Panoptic Gaussian Splatting

  • 通过查询引导的局部交叉注意力,直接从2D实例掩码生成3D分割
  • 在ScanNet-V2和ScanNet++上优于基于NeRF和高斯溅射的方法
  • 适合需要高效、准确3D全景重建的研究者与应用开发者

开放词汇全景重建是同时进行场景重建与理解的挑战性任务。近期基于高斯溅射的方法已用于3D场景理解,但多为多阶段流程,存在误差累积和对人工设计组件的依赖。为简化流程并实现全局优化,我们提出PanopticSplatting,一种端到端的开放词汇全景重建系统。该方法引入查询引导的高斯分割与局部交叉注意力机制,无需跨帧关联即可端到端提升2D实例掩码至3D。视锥内的局部交叉注意力有效降低训练内存,使模型更适用于包含大量高斯点与物体的大场景。此外,针对2D伪掩码中的噪声标签问题,提出标签混合策略以减少3D分割中的噪声漂浮物,并采用2D预测的标签投影增强多视角一致性与分割精度。实验表明,该方法在ScanNet-V2和ScanNet++数据集上显著优于基于NeRF和高斯溅射的全景重建方法。此外,PanopticSplatting可轻松推广至多种高斯溅射变体,且在不同基础模型上表现出强鲁棒性。

原文摘要 · Abstract (English)

Open-vocabulary panoptic reconstruction is a challenging task for simultaneous scene reconstruction and understanding. Recently, methods have been proposed for 3D scene understanding based on Gaussian splatting. However, these methods are multi-staged, suffering from the accumulated errors and the dependence of hand-designed components. To streamline the pipeline and achieve global optimization, we propose PanopticSplatting, an end-to-end system for open-vocabulary panoptic reconstruction. Our method introduces query-guided Gaussian segmentation with local cross attention, lifting 2D instance masks without cross-frame association in an end-to-end way. The local cross attention within view frustum effectively reduces the training memory, making our model more accessible to large scenes with more Gaussians and objects. In addition, to address the challenge of noisy labels in 2D pseudo masks, we propose label blending to promote consistent 3D segmentation with less noisy floaters, as well as label warping on 2D predictions which enhances multi-view coherence and segmentation accuracy. Our method demonstrates strong performances in 3D scene panoptic reconstruction on the ScanNet-V2 and ScanNet++ datasets, compared with both NeRF-based and Gaussian-based panoptic reconstruction methods. Moreover, PanopticSplatting can be easily generalized to numerous variants of Gaussian splatting, and we demonstrate its robustness on different Gaussian base models.

全景重建高斯溅射3D理解端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。