arXiv:2604.03309cs.CVcs.AI2026-04

用树结构引导对比学习,实现3D场景语义分割的层级一致性。

TreeGaussian: Tree-Guided Cascaded Contrastive Learning for Hierarchical Consistent 3D Gaussian Scene Segmentation and Understanding

论文配图:TreeGaussian: Tree-Guided Cascaded Contrastive Learning for Hierarchical Consistent 3D Gaussian Scene Segmentation and Understanding
图 1 · 摘自论文原文
  • 构建多层物体树,显式建模对象与部分的层次关系。
  • 两级级联对比学习,逐步优化从全局到局部的特征表示。
  • 跨视角对齐分割模式,提升复杂场景下分割质量与稳定性。

3D高斯点阵(3DGS)作为实时可微的神经场景表征方法已崭露头角,但现有基于3DGS的方法难以表达复杂的层次化3D语义结构,也难以捕捉整体-部分关系。此外,密集的成对比较和来自2D先验的不一致层级标签阻碍了特征学习,导致分割性能不佳。为此,我们提出TreeGaussian,一种树引导的级联对比学习框架,显式建模层次化语义关系并减少对比监督中的冗余。通过构建多层级物体树,实现对象-部分层次上的结构化学习。进一步提出两阶段级联对比学习策略,从全局到局部逐步精炼特征表示,缓解特征饱和并稳定训练过程。引入一致分割检测(CSD)机制与基于图的去噪模块,实现跨视角分割模式对齐,抑制不稳定高斯点,显著提升分割一致性与质量。大量实验涵盖开放词汇3D物体选择、3D点云理解及消融研究,验证了该方法的有效性与鲁棒性。

原文摘要 · Abstract (English)

3D Gaussian Splatting (3DGS) has emerged as a real-time, differentiable representation for neural scene understanding. However, existing 3DGS-based methods struggle to represent hierarchical 3D semantic structures and capture whole-part relationships in complex scenes. Moreover, dense pairwise comparisons and inconsistent hierarchical labels from 2D priors hinder feature learning, resulting in suboptimal segmentation. To address these limitations, we introduce TreeGaussian, a tree-guided cascaded contrastive learning framework that explicitly models hierarchical semantic relationships and reduces redundancy in contrastive supervision. By constructing a multi-level object tree, TreeGaussian enables structured learning across object-part hierarchies. In addition, we propose a two-stage cascaded contrastive learning strategy that progressively refines feature representations from global to local, mitigating saturation and stabilizing training. A Consistent Segmentation Detection (CSD) mechanism and a graph-based denoising module are further introduced to align segmentation modes across views while suppressing unstable Gaussian points, enhancing segmentation consistency and quality. Extensive experiments, including open-vocabulary 3D object selection, 3D point cloud understanding, and ablation studies, demonstrate the effectiveness and robustness of our approach.

3D分割层次结构对比学习高斯点阵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。