arXiv:2505.19154cs.CV2025-05NeurIPS被引 1

让3D高斯点云的语义特征更一致,提升跨视角理解能力。

FHGS: Feature-Homogenized Gaussian Splatting

  • 引入物理类比的特征融合架构,将大模型语义特征嵌入3D结构。
  • 设计非可微融合机制,实现语义特征的视角无关各向同性分布。
  • 双驱动优化提升全局语义对齐与局部结构一致性,适合3D场景理解任务。

基于3D高斯点云(3DGS)的场景理解近期取得显著进展。尽管3DGS方法具备高效渲染能力,但其高斯原语固有的各向异性颜色表示与语义特征所需的各向同性要求存在内在矛盾,导致跨视角特征一致性不足。为此,本文提出新型3D特征融合框架FHGS(Feature-Homogenized Gaussian Splatting),受物理模型启发,可在保持3DGS实时渲染效率的同时,实现预训练模型(如SAM、CLIP)任意2D特征到3D场景的高精度映射。具体创新包括:首先,提出通用特征融合架构,支持大规模预训练模型语义特征(如SAM、CLIP)在稀疏3D结构中的鲁棒嵌入;其次,引入非可微特征融合机制,使语义特征呈现视角无关的各向同性分布,从根本上调和高斯原语的各向异性渲染与特征的各向同性表达之间的矛盾;第三,提出受电势场启发的双驱动优化策略,结合语义特征场的外部监督与原语聚类的内部引导,协同优化全局语义对齐与局部结构一致性。更多交互式结果可访问:https://fhgs.cuastro.org/。

原文摘要 · Abstract (English)

Scene understanding based on 3D Gaussian Splatting (3DGS) has recently achieved notable advances. Although 3DGS related methods have efficient rendering capabilities, they fail to address the inherent contradiction between the anisotropic color representation of gaussian primitives and the isotropic requirements of semantic features, leading to insufficient cross-view feature consistency. To overcome the limitation, we proposes $\textit{FHGS}$ (Feature-Homogenized Gaussian Splatting), a novel 3D feature fusion framework inspired by physical models, which can achieve high-precision mapping of arbitrary 2D features from pre-trained models to 3D scenes while preserving the real-time rendering efficiency of 3DGS. Specifically, our $\textit{FHGS}$ introduces the following innovations: Firstly, a universal feature fusion architecture is proposed, enabling robust embedding of large-scale pre-trained models' semantic features (e.g., SAM, CLIP) into sparse 3D structures. Secondly, a non-differentiable feature fusion mechanism is introduced, which enables semantic features to exhibit viewpoint independent isotropic distributions. This fundamentally balances the anisotropic rendering of gaussian primitives and the isotropic expression of features; Thirdly, a dual-driven optimization strategy inspired by electric potential fields is proposed, which combines external supervision from semantic feature fields with internal primitive clustering guidance. This mechanism enables synergistic optimization of global semantic alignment and local structural consistency. More interactive results can be accessed on: https://fhgs.cuastro.org/.

3D高斯语义对齐特征融合实时渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。