arXiv:2503.18107cs.CV2025-03CVPR被引 23

用高斯点建模3D场景,实现开放词汇的全景分割。

PanoGS: Gaussian-based Panoptic Segmentation for 3D Open Vocabulary Scene Understanding

  • 用金字塔三平面建模连续特征空间,融合多视角2D特征云。
  • 通过语言引导图切割,将高斯点聚类为语义一致的超点。
  • 结合SAM边缘亲和度,实现高质量3D实例分割,适合室内场景理解。

最近,3D高斯泼溅(3DGS)在开放词汇场景理解任务中表现出色。然而,现有方法无法区分3D实例级信息,通常仅预测场景特征与文本查询之间的热力图。本文提出PanoGS,一种新型且高效的3D全景开放词汇场景理解方法。技术上,为学习可扩展至大型室内场景的精确3D语言特征,采用金字塔三平面建模潜在连续参数化特征空间,并使用3D特征解码器回归多视图融合的2D特征云。此外,提出语言引导图切割,协同利用重建几何与学习到的语言线索,将3D高斯原语分组为一组超原语。为获得3D一致性实例,基于图聚类分割并结合SAM引导的超原语间边缘亲和度计算。在广泛使用的数据集上进行大量实验,结果表明该方法在3D全景开放词汇场景理解任务中表现更优或更具竞争力。

原文摘要 · Abstract (English)

Recently, 3D Gaussian Splatting (3DGS) has shown encouraging performance for open vocabulary scene understanding tasks. However, previous methods cannot distinguish 3D instance-level information, which usually predicts a heatmap between the scene feature and text query. In this paper, we propose PanoGS, a novel and effective 3D panoptic open vocabulary scene understanding approach. Technically, to learn accurate 3D language features that can scale to large indoor scenarios, we adopt the pyramid tri-plane to model the latent continuous parametric feature space and use a 3D feature decoder to regress the multi-view fused 2D feature cloud. Besides, we propose language-guided graph cuts that synergistically leverage reconstructed geometry and learned language cues to group 3D Gaussian primitives into a set of super-primitives. To obtain 3D consistent instance, we perform graph clustering based segmentation with SAM-guided edge affinity computation between different super-primitives. Extensive experiments on widely used datasets show better or more competitive performance on 3D panoptic open vocabulary scene understanding. Project page: \href{https://zju3dv.github.io/panogs}{https://zju3dv.github.io/panogs}.

3D分割开放词汇高斯泼溅全景理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。