让3D场景中的语言信息精准贴合物体表面,实现更准确的语义分割与编辑。
LangSurf: Language-Embedded Surface Gaussians for 3D Scene Understanding
- 通过几何监督和对比损失,将语言特征对齐到物体表面
- 在开放词汇2D/3D分割任务上超越现有最佳方法显著表现
- 适合需要精确3D语义理解的场景编辑、移除等应用
将高斯点阵应用于3D场景理解的感知任务日益流行。现有方法多聚焦于从新视角渲染2D特征图,导致3D语言场存在误报,难以准确对齐3D空间中的物体。由于使用掩码图像提取特征,这些方法还缺乏必要上下文信息,造成特征表示不准确。为此,我们提出语言嵌入表面场(LangSurf),可精准对齐3D语言场与物体表面,支持基于文本查询的精确2D与3D分割,大幅拓展下游任务如物体移除与编辑。其核心是联合训练策略,利用几何监督和对比损失,将语言高斯投影至物体表面以赋予准确语言特征。此外,引入分层上下文感知模块,在图像级提取上下文特征,并通过SAM分割的掩码进行分层掩码池化,获得不同层级的细粒度语言特征。在开放词汇2D与3D语义分割上的大量实验表明,LangSurf显著优于先前最先进方法LangSplat。如图1所示,本方法可实现3D空间中物体的分割,显著提升实例识别、移除与编辑效果,相关实验全面验证了其有效性。
原文摘要 · Abstract (English)
Applying Gaussian Splatting to perception tasks for 3D scene understanding is becoming increasingly popular. Most existing works primarily focus on rendering 2D feature maps from novel viewpoints, which leads to an imprecise 3D language field with outlier languages, ultimately failing to align objects in 3D space. By utilizing masked images for feature extraction, these approaches also lack essential contextual information, leading to inaccurate feature representation. To this end, we propose a Language-Embedded Surface Field (LangSurf), which accurately aligns the 3D language fields with the surface of objects, facilitating precise 2D and 3D segmentation with text query, widely expanding the downstream tasks such as removal and editing. The core of LangSurf is a joint training strategy that flattens the language Gaussian on the object surfaces using geometry supervision and contrastive losses to assign accurate language features to the Gaussians of objects. In addition, we also introduce the Hierarchical-Context Awareness Module to extract features at the image level for contextual information then perform hierarchical mask pooling using masks segmented by SAM to obtain fine-grained language features in different hierarchies. Extensive experiments on open-vocabulary 2D and 3D semantic segmentation demonstrate that LangSurf outperforms the previous state-of-the-art method LangSplat by a large margin. As shown in Fig. 1, our method is capable of segmenting objects in 3D space, thus boosting the effectiveness of our approach in instance recognition, removal, and editing, which is also supported by comprehensive experiments. https://langsurf.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。