arXiv:2508.08252cs.CV2025-08ICML被引 27

根据自然语言描述,在3D场景中精准分割目标物体,支持遮挡和不可见视角下的理解。

ReferSplat: Referring Segmentation in 3D Gaussian Splatting

  • 基于空间感知的3D高斯点与语言表达显式对齐
  • 在新提出的R3DGS任务上达到领先性能
  • 适合研究多模态3D理解与具身智能的学者

我们提出3D高斯溅射中的指代分割任务(R3DGS),旨在根据自然语言描述分割3D高斯场景中的目标物体,这些描述常包含空间关系或物体属性。该任务要求模型识别在新视角下被遮挡或不可见的物体,对3D多模态理解构成重大挑战。为推动此领域研究,我们构建首个R3DGS数据集Ref-LERF。分析表明,3D多模态理解与空间关系建模是关键难点。为此,我们提出ReferSplat框架,通过空间感知方式显式建模3D高斯点与自然语言表达的关系。该方法在新提出的R3DGS任务及3D开放词汇分割基准上均取得当前最优表现。代码与数据集已公开于https://github.com/heshuting555/ReferSplat。

原文摘要 · Abstract (English)

We introduce Referring 3D Gaussian Splatting Segmentation (R3DGS), a new task that aims to segment target objects in a 3D Gaussian scene based on natural language descriptions, which often contain spatial relationships or object attributes. This task requires the model to identify newly described objects that may be occluded or not directly visible in a novel view, posing a significant challenge for 3D multi-modal understanding. Developing this capability is crucial for advancing embodied AI. To support research in this area, we construct the first R3DGS dataset, Ref-LERF. Our analysis reveals that 3D multi-modal understanding and spatial relationship modeling are key challenges for R3DGS. To address these challenges, we propose ReferSplat, a framework that explicitly models 3D Gaussian points with natural language expressions in a spatially aware paradigm. ReferSplat achieves state-of-the-art performance on both the newly proposed R3DGS task and 3D open-vocabulary segmentation benchmarks. Dataset and code are available at https://github.com/heshuting555/ReferSplat.

3D分割多模态语言理解高斯溅射

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。