用稀疏无姿态图像实现快速3D场景重建与语义对齐
LangFlash: Feed-forward 3D Language Gaussian Splatting from Sparse Unposed Images

- 单次前向传播直接预测几何与语义,无需迭代优化
- 在RealEstate10k上构建稠密语义标注,支持大规模训练
- 稀疏语义编码保留语言信息,降低表示复杂度
我们提出LangFlash,一种前向型3D语言高斯点云重建框架,可从稀疏无姿态多视角图像中,以高斯原语参数化场景,并注入语义对齐特征。不同于依赖优化的3D方法,LangFlash通过一次前向传播直接输出几何与语义,实现低延迟重建和语言一致的场景理解。为支持大规模训练,我们在RealEstate10k数据集上引入连贯且稠密的3D语义标注,提供语义监督。此外,提出一种稀疏语义编码方案,结合全局语义词典与局部逐原语权重,既保留高层语言信息,又降低表示复杂度。实验表明,相比先前方法,LangFlash在新视角合成与语义一致性上表现更优。本研究建立了一种无姿态、语言引导的3D场景重建新范式,推动通用3D视觉与多模态场景理解的发展。演示见https://liylo.github.io/langflash.github.io/。
原文摘要 · Abstract (English)
We present LangFlash, a feed-forward framework for 3D Language Gaussian Splatting that reconstructs 3D scenes parameterized by Gaussian primitives enriched with language-aligned semantic features from sparse unposed multi-view images. Unlike optimization-based 3D methods, LangFlash directly predicts the geometry and semantics in a single forward pass, enabling low-latency 3D reconstruction and language-consistent scene understanding. To support large-scale training, we enriched the RealEstate10k dataset with coherent and dense semantic information for 3D semantic supervision. Furthermore, we propose a sparse semantic encoding scheme that combines a global semantic dictionary with locally varying per-primitive weights, preserving high-level linguistic information, while reducing representation complexity. Experimental results show that LangFlash achieves superior novel view synthesis and semantic consistency compared with previous methods. This study establishes a new paradigm for pose-free, language-grounded 3D scene reconstruction, advancing generalizable 3D vision and multimodal scene understanding. Demo is available at https://liylo.github.io/langflash.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。