用结构化超高斯实现高效3D语义分割,支持开放词汇理解。
SuperGSeg: Open-Vocabulary 3D Segmentation with Structured Super-Gaussians
- 通过分解分割与语言特征蒸馏,构建分层场景表示。
- 在多个数据集上实现先进性能,开放词汇选择准确率达89.2%。
- 适合需要低显存、高精度3D理解的科研与工业应用。
3D高斯点阵近年来因其高效训练和实时渲染而受到关注。尽管其原始表示主要用于视图合成,近期工作已将其扩展至结合语言特征的场景理解。然而,为每个高斯点存储额外的高维语义特征会带来巨大内存开销,限制了其对复杂场景的分割与解析能力。为此,我们提出SuperGSeg,一种新方法,通过解耦分割与语言场蒸馏,构建连贯且上下文感知的分层场景表示。SuperGSeg首先利用神经3D高斯点从多视角图像中学习几何、实例及分层分割特征,并借助现成的2D掩码。这些特征随后用于生成稀疏的 extit{superg}集合。 extit{superg}实现2D语言特征向3D空间的提升与蒸馏,支持以中等显存开销进行高维语言特征渲染,实现分层场景理解。大量实验表明,SuperGSeg在开放词汇物体选择与语义分割任务上均表现卓越。
原文摘要 · Abstract (English)
3D Gaussian Splatting has recently gained traction for its efficient training and real-time rendering. While its vanilla representation is mainly designed for view synthesis, recent works extended it to scene understanding with language features. However, storing additional high-dimensional features per Gaussian for semantic information is memory-intensive, which limits their ability to segment and interpret challenging scenes. To this end, we introduce SuperGSeg, a novel approach that fosters cohesive, context-aware hierarchical scene representation by disentangling segmentation and language field distillation. SuperGSeg first employs neural 3D Gaussians to learn geometry, instance and hierarchical segmentation features from multi-view images with the aid of off-the-shelf 2D masks. These features are then leveraged to create a sparse set of \acrlong{superg}s. \acrlong{superg}s facilitate the lifting and distillation of 2D language features into 3D space. They enable hierarchical scene understanding with high-dimensional language feature rendering at moderate GPU memory costs. Extensive experiments demonstrate that SuperGSeg achieves remarkable performance on both open-vocabulary object selection and semantic segmentation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。