arXiv:2608.10057cs.CV2026-08中稿 · ECCV

将2D语义分割升级为3D层级化场景理解,支持跨视角一致的语义推理。

LEGO: Leveled Language Gaussian Splatting

论文配图:LEGO: Leveled Language Gaussian Splatting
图 1 · 摘自论文原文
  • 基于SAM的多粒度2D分割,自适应融合为统一3D语义层级结构。
  • 在ScanNet和Matterport3D上实现新SOTA,支持开放词汇3D分割。
  • 构建分层语言场景图,助力大模型进行空间关系推理与精准定位。

我们提出LEGO,用于高级开放词汇场景理解。其核心创新在于捕捉场景中的内在语义层级,如'花盆→花束→花苞→花瓣'的演化关系。尽管基础模型如SAM可在2D中识别多粒度结构,但其分割严格依赖视角且缺乏跨视图一致性。LEGO将多视图中不稳定的SAM粒度自适应地重分级为统一、3D一致的层级结构,为3D场景提供精确的结构化多级分割监督。通过使用CLIP嵌入对这些片段进行语义锚定,LEGO在层级间恢复开放词汇语义逻辑。进一步结合空间关系,将这些片段提升为分层语言场景图,有效赋能大语言模型进行复杂、上下文感知的空间推理与精确视觉定位。实验表明,LEGO在可提示与开放词汇3D分割基准上均达到新SOTA,展现出先进的层次化场景分解与上下文感知空间推理能力。

原文摘要 · Abstract (English)

We introduce LEGO for advanced open-vocabulary scene understanding. Beyond basic concept recognition, its core innovation lies in capturing the intrinsic semantic hierarchies within the scene, such as the "flowerpot -> bouquet -> bud -> petal" lineage. While foundation models like SAM can identify multi-granular structures in 2D, their partitions are strictly perspective-bound and lack cross-view consensus. LEGO self-adaptively re-grades volatile multi-view SAM granularities into a unified, 3D-consistent hierarchy. This provides precise supervision for the structurally coherent, multi-level segmentation of 3D scenes. By grounding these segments with CLIP embeddings, LEGO recovers open-vocabulary semantic logic across hierarchical levels. Furthermore, by incorporating spatial relationships, we elevate these segments into level-wise language scene graphs, effectively empowering Large Language Models to perform complex, context-aware spatial reasoning and precise visual grounding. Experimental results demonstrate that LEGO establishes new state-of-the-art performance across both promptable and open-vocabulary 3D segmentation benchmarks, exhibiting advanced hierarchical scene decomposition and context-aware spatial reasoning.

3D理解语义层级语言场景图开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。