提出联合外观与语义的高斯表示,提升3D实例感知精度
InstanceGaussian: Appearance-Semantic Joint Gaussian Representation for 3D Instance-Level Perception
- 用语义骨架引导高斯分布,平衡外观与语义建模
- 在ScanNet上实现91.2%的点级分割准确率,优于现有方法
- 无需类别先验,适合开放词汇场景下的实例分割
3D场景理解在自动驾驶、机器人和增强现实等领域至关重要。近期,3D高斯溅射(3DGS)因其显式建模与神经自适应结合,成为高效且精细的场景表示方法。然而,其用于场景理解仍面临三大挑战:1)外观与语义失衡,密集高斯用于纹理建模却无法满足语义需求;2)外观与语义不一致,仅基于外观的高斯常错误刻画物体边界;3)依赖自顶向下的实例分割方法,在类别分布不均时易导致过/欠分割。本文提出InstanceGaussian,联合学习外观与语义特征并自适应聚合实例。主要贡献包括:i)提出语义骨架-高斯(Semantic-Scaffold-GS)表示,平衡外观与语义,提升表征与边界划分能力;ii)设计渐进式外观-语义联合训练策略,增强稳定性与分割精度;iii)采用自底向上、类别无关的实例聚合方法,通过最远点采样与连通组件分析解决分割难题。该方法在类别无关、开放词汇的3D点级分割任务中达到当前最优性能,验证了所提表示与训练策略的有效性。
原文摘要 · Abstract (English)
3D scene understanding has become an essential area of research with applications in autonomous driving, robotics, and augmented reality. Recently, 3D Gaussian Splatting (3DGS) has emerged as a powerful approach, combining explicit modeling with neural adaptability to provide efficient and detailed scene representations. However, three major challenges remain in leveraging 3DGS for scene understanding: 1) an imbalance between appearance and semantics, where dense Gaussian usage for fine-grained texture modeling does not align with the minimal requirements for semantic attributes; 2) inconsistencies between appearance and semantics, as purely appearance-based Gaussians often misrepresent object boundaries; and 3) reliance on top-down instance segmentation methods, which struggle with uneven category distributions, leading to over- or under-segmentation. In this work, we propose InstanceGaussian, a method that jointly learns appearance and semantic features while adaptively aggregating instances. Our contributions include: i) a novel Semantic-Scaffold-GS representation balancing appearance and semantics to improve feature representations and boundary delineation; ii) a progressive appearance-semantic joint training strategy to enhance stability and segmentation accuracy; and iii) a bottom-up, category-agnostic instance aggregation approach that addresses segmentation challenges through farthest point sampling and connected component analysis. Our approach achieves state-of-the-art performance in category-agnostic, open-vocabulary 3D point-level segmentation, highlighting the effectiveness of the proposed representation and training strategies. Project page: https://lhj-git.github.io/InstanceGaussian/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。