用代码本注意力增强3D高斯场的语义一致性,实现开放词汇场景理解。
OpenGaFF: Open-Vocabulary Gaussian Feature Field with Codebook Attention

- 构建连续语义函数,将语义与几何结构显式耦合。
- 在多个基准上超越现有方法,提升3D语义一致性和分割质量。
- 适合做3D开放词汇理解、语义可解释性研究的开发者。
基于高斯表示的开放词汇3D场景理解仍面临多视角观测下语义预测碎片化与空间不一致的问题。本文提出OpenGaFF,一种基于3D高斯点云渲染的新型框架。核心是高斯特征场,将语义建模为高斯几何与外观的连续函数,通过显式依赖几何结构强化语义与几何的耦合,提升相似结构在3D空间中的空间一致性。为进一步保障物体级语义一致性,引入结构化代码本作为共享语义基元,并设计代码本引导的注意力机制,通过查询嵌入与代码本条目间的相似性匹配检索语言特征,实现鲁棒的开放词汇推理并降低物体内部特征方差。在标准2D与3D开放词汇基准上的大量实验表明,该方法持续优于先前方法,实现了更优的分割质量、更强的3D语义一致性以及语义可解释的代码本,揭示了学习表征的内在机制。
原文摘要 · Abstract (English)
Understanding open-vocabulary 3D scenes with Gaussian-based representations remains challenging due to fragmented and spatially inconsistent semantic predictions across multi-view observations. In this paper, we present OpenGaFF, a novel framework for open-vocabulary 3D scene understanding built upon 3D Gaussian Splatting. At the core of our method is a Gaussian Feature Field that models semantics as a continuous function of Gaussian geometry and appearance. By explicitly conditioning semantic predictions on geometric structure, this formulation strengthens the coupling between geometry and semantics, leading to improved spatial coherence across similar structures in 3D space. To further enforce object-level semantic consistency, we introduce a structured codebook that serves as a set of shared semantic primitives. Furthermore, a codebook-guided attention mechanism is proposed to retrieve language features via similarity matching between query embeddings and learned codebook entries, enabling robust open-vocabulary reasoning while reducing intra-object feature variance. Extensive experiments on standard 2D and 3D open-vocabulary benchmarks demonstrate that our method consistently outperforms prior approaches, achieving improved segmentation quality, stronger 3D semantic consistency and a semantically interpretable codebook that provides insight into the learned representation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。