arXiv:2506.13030cs.CV2025-06NeurIPS被引 2

用野外多样图片训练3D视角生成,让模型自适应不同光照和遮挡。

WildCAT3D: Appearance-Aware Multi-View Diffusion in the Wild

  • 通过建模全局外观条件,让扩散模型从杂乱野外图像中学习多视角3D结构。
  • 在物体与场景级单视图重建上达到当前最佳效果,训练数据更少。
  • 支持生成时控制整体光照风格,适合需要外观一致性的应用。

尽管稀疏新视角合成(NVS)在以物体为中心的场景中取得进展,但场景级NVS仍是难题。核心瓶颈在于缺乏可用的干净多视角训练数据,现有手动标注数据集多样性、相机变化或授权问题受限。另一方面,大量多样且可自由使用的野外数据存在,如旅游照片,包含不同光照、瞬时遮挡等外观差异。为此,我们提出WildCAT3D,一种从野外采集的多样化2D场景图像中学习并生成新视角的框架。通过显式建模图像中的全局外观条件,将最先进的多视角扩散范式扩展至可学习具有外观差异的场景视图。训练后的模型在推理时能泛化到新场景,生成多个一致的新视角。WildCAT3D在物体级与场景级单视图NVS任务中均达到当前最佳性能,且训练所用数据源更少。此外,生成过程中提供全局外观控制,拓展了新应用场景。

原文摘要 · Abstract (English)

Despite recent advances in sparse novel view synthesis (NVS) applied to object-centric scenes, scene-level NVS remains a challenge. A central issue is the lack of available clean multi-view training data, beyond manually curated datasets with limited diversity, camera variation, or licensing issues. On the other hand, an abundance of diverse and permissively-licensed data exists in the wild, consisting of scenes with varying appearances (illuminations, transient occlusions, etc.) from sources such as tourist photos. To this end, we present WildCAT3D, a framework for generating novel views of scenes learned from diverse 2D scene image data captured in the wild. We unlock training on these data sources by explicitly modeling global appearance conditions in images, extending the state-of-the-art multi-view diffusion paradigm to learn from scene views of varying appearances. Our trained model generalizes to new scenes at inference time, enabling the generation of multiple consistent novel views. WildCAT3D provides state-of-the-art results on single-view NVS in object- and scene-level settings, while training on strictly less data sources than prior methods. Additionally, it enables novel applications by providing global appearance control during generation.

3D生成扩散模型多视角合成外观控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。