arXiv:2503.18052cs.CV2025-03ICCV被引 49

首个直接在3D高斯点云上做语义理解的模型,无需2D或文本辅助。

SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language Pretraining

  • 基于3D高斯点云原生训练,实现端到端3D语义学习
  • 在7916个室内场景数据集上显著超越现有基线方法
  • 适合需要纯3D理解的机器人、AR/VR场景应用

识别任意或未见类别是实现完整真实世界3D场景理解的关键。当前所有方法均依赖2D图像或文本模态进行训练,或在推理时结合使用,缺乏仅处理3D数据并实现端到端语义学习的模型及相应数据。3D高斯点云(3DGS)已成为多种视觉任务中3D场景表示的标准。然而,如何通用地将语义推理融入3DGS仍是开放挑战。为此,我们提出SceneSplat,据我们所知首个原生基于3DGS的大型室内场景理解方法。同时提出一种自监督学习方案,从无标注场景中挖掘丰富的3D特征。为支持该方法,我们构建了SceneSplat-7K——首个面向室内场景的大型3DGS数据集,包含7916个来自ScanNet、Matterport3D等七个基准数据集的场景。生成该数据集耗时相当于150个L4 GPU天,为3DGS-based推理提供标准化基准。在SceneSplat-7K上的大量实验表明,该方法显著优于现有基线。

原文摘要 · Abstract (English)

Recognizing arbitrary or previously unseen categories is essential for comprehensive real-world 3D scene understanding. Currently, all existing methods rely on 2D or textual modalities during training or together at inference. This highlights the clear absence of a model capable of processing 3D data alone for learning semantics end-to-end, along with the necessary data to train such a model. Meanwhile, 3D Gaussian Splatting (3DGS) has emerged as the de facto standard for 3D scene representation across various vision tasks. However, effectively integrating semantic reasoning into 3DGS in a generalizable manner remains an open challenge. To address these limitations, we introduce SceneSplat, to our knowledge the first large-scale 3D indoor scene understanding approach that operates natively on 3DGS. Furthermore, we propose a self-supervised learning scheme that unlocks rich 3D feature learning from unlabeled scenes. To power the proposed methods, we introduce SceneSplat-7K, the first large-scale 3DGS dataset for indoor scenes, comprising 7916 scenes derived from seven established datasets, such as ScanNet and Matterport3D. Generating SceneSplat-7K required computational resources equivalent to 150 GPU days on an L4 GPU, enabling standardized benchmarking for 3DGS-based reasoning for indoor scenes. Our exhaustive experiments on SceneSplat-7K demonstrate the significant benefit of the proposed method over the established baselines.

3D理解高斯点云自监督学习场景理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。