提出高效且多视角一致的开放词汇3D语义高斯模型,提升分割精度与计算效率。
econSG: Efficient and Multi-view Consistent Open-Vocabulary 3D Semantic Gaussians
- 通过双向优化SAM与CLIP,实现精确语义边界。
- 融合多视角特征后统一降维,保持3D一致性并提速。
- 在4个数据集上达顶尖性能,训练速度最快。
近期开放词汇神经场研究主要聚焦于从视觉语言模型(VLMs)提取精准语义特征,并高效构建多视角一致的3D神经场表示。然而,多数方法过度依赖SAM对图像级CLIP特征进行正则化,缺乏进一步优化。此外,部分方法在融合前对2D VLM语义特征降维,导致多视角不一致。本文提出econSG,用于开放词汇3D语义分割。其包含:1)置信区域引导正则化(CRR),双向精炼SAM与CLIP,获得兼具完整性和精确边界的语义特征;2)低维上下文空间,在融合反投影多视角2D特征后,直接对3D特征进行降维,而非分别处理各2D视图,从而增强3D多视角一致性并提升效率。econSG在四个基准数据集上表现优于现有方法,且为所有方法中训练最高效。
原文摘要 · Abstract (English)
The primary focus of most recent works on open-vocabulary neural fields is extracting precise semantic features from the VLMs and then consolidating them efficiently into a multi-view consistent 3D neural fields representation. However, most existing works over-trusted SAM to regularize image-level CLIP without any further refinement. Moreover, several existing works improved efficiency by dimensionality reduction of semantic features from 2D VLMs before fusing with 3DGS semantic fields, which inevitably leads to multi-view inconsistency. In this work, we propose econSG for open-vocabulary semantic segmentation with 3DGS. Our econSG consists of: 1) A Confidence-region Guided Regularization (CRR) that mutually refines SAM and CLIP to get the best of both worlds for precise semantic features with complete and precise boundaries. 2) A low dimensional contextual space to enforce 3D multi-view consistency while improving computational efficiency by fusing backprojected multi-view 2D features and follow by dimensional reduction directly on the fused 3D features instead of operating on each 2D view separately. Our econSG shows state-of-the-art performance on four benchmark datasets compared to the existing methods. Furthermore, we are also the most efficient training among all the methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。