用无约束照片生成可交互的3D建筑语义场景,支持开放式查询。
Taking Language Embedded 3D Gaussian Splatting into the Wild
- 基于多视角渲染与CLIP特征,构建语言嵌入的3D高斯点云
- 在PT-OVS数据集上开集分割精度超越现有方法
- 适合做建筑风格分析、3D编辑和开放词汇查询的开发者
近期利用大规模互联网照片进行3D重建的技术,已实现全球地标与历史遗址的沉浸式虚拟探索。然而,对建筑风格与结构知识的沉浸式理解仍局限于静态图文浏览。本文提出一种新框架,从无约束照片集合中实现开放词汇的3D场景理解。首先,从同一视角渲染多张外观图像,提取多外观CLIP特征,并生成瞬态与外观不确定性图以指导优化;接着,设计瞬态不确定性感知自编码器、多外观语言场3DGS表示及后融合策略,有效压缩、学习并融合多视角语言特征;最后,引入新基准数据集PT-OVS,用于评估开放词汇分割性能。实验表明,该方法显著优于现有方法,在准确率与应用灵活性上均有提升,支持交互漫游、建筑风格识别与3D场景编辑等任务。
原文摘要 · Abstract (English)
Recent advances in leveraging large-scale Internet photo collections for 3D reconstruction have enabled immersive virtual exploration of landmarks and historic sites worldwide. However, little attention has been given to the immersive understanding of architectural styles and structural knowledge, which remains largely confined to browsing static text-image pairs. Therefore, can we draw inspiration from 3D in-the-wild reconstruction techniques and use unconstrained photo collections to create an immersive approach for understanding the 3D structure of architectural components? To this end, we extend language embedded 3D Gaussian splatting (3DGS) and propose a novel framework for open-vocabulary scene understanding from unconstrained photo collections. Specifically, we first render multiple appearance images from the same viewpoint as the unconstrained image with the reconstructed radiance field, then extract multi-appearance CLIP features and two types of language feature uncertainty maps-transient and appearance uncertainty-derived from the multi-appearance features to guide the subsequent optimization process. Next, we propose a transient uncertainty-aware autoencoder, a multi-appearance language field 3DGS representation, and a post-ensemble strategy to effectively compress, learn, and fuse language features from multiple appearances. Finally, to quantitatively evaluate our method, we introduce PT-OVS, a new benchmark dataset for assessing open-vocabulary segmentation performance on unconstrained photo collections. Experimental results show that our method outperforms existing methods, delivering accurate open-vocabulary segmentation and enabling applications such as interactive roaming with open-vocabulary queries, architectural style pattern recognition, and 3D scene editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。