arXiv:2503.21767cs.CV2025-03被引 5

让3D场景按语言指令精准定位,提升机器人理解能力。

Semantic Consistent Language Gaussian Splatting for Point-Level Open-vocabulary Querying

  • 用分割掩码追踪建立语义一致的标签,指导语言高斯分布训练。
  • 先找整体语义真值,再定位具体高斯点,查询准确率显著提升。
  • 适合需要语言控制3D场景的机器人应用,如导航与操作。

开放词汇3D场景理解对机器人应用至关重要,如自然语言驱动的操作、人机交互和自主导航。现有基于3D高斯溅射的查询方法常因2D掩码监督不一致且缺乏鲁棒的3D点级检索机制而受限。本文提出一种新型点级查询框架:(i)通过在分割掩码上进行追踪,建立语义一致的真值,用于蒸馏语言高斯;(ii)引入基于真值锚定的查询方法,先检索蒸馏后的真值,再用其查询单个高斯点。在三个基准数据集上的大量实验表明,该方法优于现有最先进水平。在LERF、3D-OVS和Replica数据集上,mIoU分别提升+4.14、+20.42和+1.7。结果验证了该框架在真实机器人系统中实现开放词汇理解的潜力。

原文摘要 · Abstract (English)

Open-vocabulary 3D scene understanding is crucial for robotics applications, such as natural language-driven manipulation, human-robot interaction, and autonomous navigation. Existing methods for querying 3D Gaussian Splatting often struggle with inconsistent 2D mask supervision and lack a robust 3D point-level retrieval mechanism. In this work, (i) we present a novel point-level querying framework that performs tracking on segmentation masks to establish a semantically consistent ground-truth for distilling the language Gaussians; (ii) we introduce a GT-anchored querying approach that first retrieves the distilled ground-truth and subsequently uses the ground-truth to query the individual Gaussians. Extensive experiments on three benchmark datasets demonstrate that the proposed method outperforms state-of-the-art performance. Our method achieves an mIoU improvement of +4.14, +20.42, and +1.7 on the LERF, 3D-OVS, and Replica datasets. These results validate our framework as a promising step toward open-vocabulary understanding in real-world robotic systems.

3D生成语言理解机器人高斯溅射

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。