arXiv:2506.09565cs.CV2025-06被引 14

用语言感知的高斯场实现高效3D场景语义重建

SemanticSplat: Feed-Forward 3D Scene Understanding with Language-Aware Gaussian Fields

  • 将3D高斯与潜在语义属性融合,统一建模几何、外观和语义
  • 仅需稀疏视角图像即可生成高质量语义3D场景,无优化步骤
  • 支持提示驱动和开放词汇分割,适合增强现实与机器人应用

完整3D场景理解需联合建模几何、外观与语义,对增强现实与机器人交互至关重要。现有前馈式方法(如LSM)仅能提取语言语义,难以实现整体理解,且几何重建质量差、存在噪声。而逐场景优化方法依赖密集输入视图,部署复杂。本文提出SemanticSplat,一种前馈式语义感知3D重建方法,将3D高斯与潜在语义属性统一建模。通过融合多特征场(如LSeg、SAM)与存储跨视图特征相似性的代价体表示,生成语义各向异性高斯。采用两阶段蒸馏框架,从稀疏视图图像重建出多模态语义特征场。实验表明该方法在可提示与开放词汇分割等任务上表现优异。视频演示见https://semanticsplat.github.io。

原文摘要 · Abstract (English)

Holistic 3D scene understanding, which jointly models geometry, appearance, and semantics, is crucial for applications like augmented reality and robotic interaction. Existing feed-forward 3D scene understanding methods (e.g., LSM) are limited to extracting language-based semantics from scenes, failing to achieve holistic scene comprehension. Additionally, they suffer from low-quality geometry reconstruction and noisy artifacts. In contrast, per-scene optimization methods rely on dense input views, which reduces practicality and increases complexity during deployment. In this paper, we propose SemanticSplat, a feed-forward semantic-aware 3D reconstruction method, which unifies 3D Gaussians with latent semantic attributes for joint geometry-appearance-semantics modeling. To predict the semantic anisotropic Gaussians, SemanticSplat fuses diverse feature fields (e.g., LSeg, SAM) with a cost volume representation that stores cross-view feature similarities, enhancing coherent and accurate scene comprehension. Leveraging a two-stage distillation framework, SemanticSplat reconstructs a holistic multi-modal semantic feature field from sparse-view images. Experiments demonstrate the effectiveness of our method for 3D scene understanding tasks like promptable and open-vocabulary segmentation. Video results are available at https://semanticsplat.github.io.

3D重建语义理解高斯场多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。