仅用稀疏图像重建可通用的3D语言场景,避免渲染伪影。
LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion
- 用TriMap扩散模型从稀疏视图生成外观、几何与语义信息。
- 通过语言压缩器实现跨场景泛化,无需逐场景重训练。
- 支持开放词汇查询,适合需要多模态3D理解的应用。
从2D图像中恢复具有开放词汇语义理解的3D结构是一项基础但极具挑战的任务。现有方法依赖校准的密集视角重建范式,导致在视图有限时出现严重渲染伪影和不合理的语义合成。本文提出LangScene-X,一种新型生成框架,统一生成3D一致的多模态信息以实现重建与理解。利用生成模型创造更一致的新观测的能力,我们仅需稀疏视图即可构建可泛化的3D语言嵌入场景。具体而言,首先训练一个TriMap视频扩散模型,通过渐进式知识融合,从稀疏输入生成外观(RGB)、几何(法向量)和语义(分割图)。此外,提出语言量化压缩器(LQC),基于大规模图像数据集训练,高效编码语言嵌入,实现跨场景泛化且无需逐场景重训练。最后,通过将语言信息对齐到3D场景表面,重构语言表面场,支持开放式语言查询。在真实世界数据上的大量实验表明,与现有最优方法相比,LangScene-X在质量与泛化性方面均具优势。
原文摘要 · Abstract (English)
Recovering 3D structures with open-vocabulary scene understanding from 2D images is a fundamental but daunting task. Recent developments have achieved this by performing per-scene optimization with embedded language information. However, they heavily rely on the calibrated dense-view reconstruction paradigm, thereby suffering from severe rendering artifacts and implausible semantic synthesis when limited views are available. In this paper, we introduce a novel generative framework, coined LangScene-X, to unify and generate 3D consistent multi-modality information for reconstruction and understanding. Powered by the generative capability of creating more consistent novel observations, we can build generalizable 3D language-embedded scenes from only sparse views. Specifically, we first train a TriMap video diffusion model that can generate appearance (RGBs), geometry (normals), and semantics (segmentation maps) from sparse inputs through progressive knowledge integration. Furthermore, we propose a Language Quantized Compressor (LQC), trained on large-scale image datasets, to efficiently encode language embeddings, enabling cross-scene generalization without per-scene retraining. Finally, we reconstruct the language surface fields by aligning language information onto the surface of 3D scenes, enabling open-ended language queries. Extensive experiments on real-world data demonstrate the superiority of our LangScene-X over state-of-the-art methods in terms of quality and generalizability. Project Page: https://liuff19.github.io/LangScene-X.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。