无需标定与优化,用稀疏图像对实现通用3D语义建模
GSemSplat: Generalizable Semantic 3D Gaussian Splatting from Uncalibrated Image Pairs
- 基于双特征监督,从非标定图像对学习3D高斯的开放词汇语义
- 在ScanNet++上优于传统方法,实现端到端泛化3D语义建模
- 适合需要快速部署的AR/机器人场景,无需每场景重训练
三维世界建模与理解对增强现实、机器人导航等应用至关重要。现有基于3D高斯点阵的方法虽能融合多视角图像的语义信息,但通常依赖密集校准图像的逐场景优化,实用性受限。本文提出新任务:从稀疏非标定图像对中实现可泛化的3D语义场建模。基于Splatt3R架构,我们设计了GSemSplat框架,无需逐场景优化、密集图像或相机标定,即可将开放词汇语义映射至3D高斯。通过在2D空间使用区域特异性与上下文感知语义特征双重监督,有效互补提升3D语义学习效果。在ScanNet++数据集上的实验表明,该方法显著优于传统场景特定方法,展现了优越性与有效性。希望推动通用3D理解的研究发展。
原文摘要 · Abstract (English)
Modeling and understanding the 3D world is crucial for various applications, from augmented reality to robotic navigation. Recent advancements based on 3D Gaussian Splatting have integrated semantic information from multi-view images into Gaussian primitives. However, these methods typically require costly per-scene optimization from dense calibrated images, limiting their practicality. In this paper, we consider the new task of generalizable 3D semantic field modeling from sparse, uncalibrated image pairs. Building upon the Splatt3R architecture, we introduce GSemSplat, a framework that learns open-vocabulary semantic representations linked to 3D Gaussians without the need for per-scene optimization, dense image collections or calibration. To ensure effective and reliable learning of semantic features in 3D space, we employ a dual-feature approach that leverages both region-specific and context-aware semantic features as supervision in the 2D space. This allows us to capitalize on their complementary strengths. Experimental results on the ScanNet++ dataset demonstrate the effectiveness and superiority of our approach compared to the traditional scene-specific method. We hope our work will inspire more research into generalizable 3D understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。