arXiv:2606.11894cs.CV2026-06

用海量真实照片训练3D高斯点云,无需逐场景优化也能还原复杂光照下的场景。

Wild3R: Feed-Forward 3D Gaussian Splatting from Unconstrained Sparse Photo Collection

论文配图:Wild3R: Feed-Forward 3D Gaussian Splatting from Unconstrained Sparse Photo Collection
图 1 · 摘自论文原文
  • 基于参考视图学习跨视角外观一致性,自动剔除临时物体
  • 在33.7万张图像的WildCity数据集上训练,覆盖170种光照条件
  • 适用于无约束的真实照片集合,性能媲美需优化的旧方法

前馈式3D高斯点云(3DGS)无需对每个场景进行耗时优化,但现有方法在包含多样光照和临时物体的真实照片中表现不佳。本文提出Wild3R,一种面向无约束稀疏照片集的前馈方法。核心瓶颈在于缺乏同时具备多视角、多样光照和动态变化的训练数据。为此,我们构建了WildCity数据集,包含200个场景、170种照明条件和动态物体,总计337,500张图像。利用该数据集,模型学习在参考视图条件下保持跨视角外观一致性,并去除瞬态内容。大量实验表明,本方法优于现有前馈方法,且性能可与依赖场景优化的先前方法相媲美。

原文摘要 · Abstract (English)

Feed-forward 3D Gaussian Splatting (3DGS) removes the need for time-consuming per-scene optimization required by traditional 3DGS. However, existing feed-forward approaches struggle with real-world photo collections that include diverse lighting conditions and transient objects. In this paper, we present Wild3R, a feed-forward approach for unconstrained sparse photo collections. The main bottleneck is the lack of training data that provides multiple viewpoints, a variety of illuminations, and transient variations necessary for learning robust scene representations. To address this, we introduce the WildCity dataset, which comprises 200 scenes, 170 lighting conditions, and transient objects, resulting in 337,500 images in total. By leveraging the dataset, our model learns appearance consistency across viewpoints conditioned on reference views, while removing transient content. Extensive experiments demonstrate that our method outperforms existing feed-forward approaches and achieves results competitive with prior per-scene optimization-based methods.

3D重建高斯点云前馈生成真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。