让无姿态野外照片也能快速生成新视角,还能随意换光影风格。
WildSplat: Feedforward Gaussian Splatting from Unposed In-the-Wild Images

- 双分支结构分离几何与外观,分别处理场景结构和光照变化。
- 单次前向传播实现高质量新视角合成,优于现有优化和前馈方法。
- 适合做野外照片的3D重建与外观编辑,尤其擅长处理复杂光照。
尽管前馈式3D重建在新视角合成中效率高,但在光照变化环境下表现不佳。为此,我们提出WildSplat,首个可对无姿态野外图像进行外观条件化新视角合成的前馈3D高斯点云框架。为应对不一致的光照条件,我们设计了双分支架构,显式分离几何与外观:几何分支提取与外观无关的3D结构,并联合预测相机位姿;外观分支通过全局预调制交叉注意力机制,将目标外观特征注入内容特征中以控制渲染效果。为进一步防止特征混淆,引入联合多参考训练策略,稳定训练过程。大量实验表明,WildSplat在稀疏输入条件下,单次前向传播即可实现领先的野外新视角合成与外观编辑性能,超越现有基于优化及前馈的方法。
原文摘要 · Abstract (English)
While feedforward 3D reconstruction excels at efficient novel view synthesis, it typically falters when faced with scenes under varying illumination. To this end, we introduce WildSplat, the first feedforward 3D Gaussian Splatting framework capable of appearance-conditioned novel-view synthesis for unposed in-the-wild images. To handle inconsistent photometric conditions, we propose a dual-branch architecture that explicitly decouples geometry from appearance. The geometry branch extracts an appearance-invariant 3D structure and jointly predicts camera poses. To govern the rendering appearance, the appearance branch injects target appearance cues into the content features via a globally pre-modulated cross-attention mechanism. To further prevent feature entanglement, we introduce a joint multi-reference training strategy that stabilizes the training process. Extensive experiments show that WildSplat surpasses existing optimization-based and feedforward methods, achieving state-of-the-art performance in in-the-wild novel view synthesis and appearance editing from sparse inputs in a single forward pass.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。