无需训练,从单图或视频生成高质量新视角画面
NeoMap: Training-free Novel-View Synthesis from Single Images and Videos

- 利用预训练模型中隐含的视频数据流形,通过迭代优化噪声来定位最佳新视角解
- 在Tanks-and-Temples、LLFF、DAVIS等3个基准上均超越现有方法,图像保真度和视角一致性领先
- 适用于希望快速部署新视角生成而无需微调的开发者或研究者
我们研究从单张图像或单目视频中进行新视角视频合成这一挑战性问题。现有方法通常假设预训练视频模型不具备原生新视角合成能力,需依赖相机条件、任务特定微调或分步硬去噪引导来实现视角对齐,但常导致伪影和全局场景不一致。本文提出NeoMap,一种无需训练的新框架,旨在从通用预训练视频模型中找到高保真、视角一致的新视角解。核心思想是:理想的视角解已自然编码于预训练模型所学习的视频数据流形中,关键在于精准定位该最优解。我们通过核心机制——收敛的流形交替投影迭代,优化初始噪声来实现。大量实验表明,NeoMap在三个标准新视角合成基准(包括具有挑战性的Tanks-and-Temples、LLFF和DAVIS数据集)上显著优于所有现有方法,实现了最先进的生成保真度与顶级的视角一致性。
原文摘要 · Abstract (English)
We study the challenging problem of novel view video synthesis from single images or monocular videos. Existing methods, which operate under the assumption that pre-trained video models lack native novel view synthesis capability and enforce view alignment via camera conditioning, task-specific fine-tuning, or stepwise hard denoising guidance, often suffer from artifacts and compromised global scene consistency. In this paper, we introduce NeoMap, a novel training-free framework designed to locate high-fidelity, view-consistent novel view solutions from general pre-trained video models. The key to our approach is the core insight that promising novel view solutions are inherently encoded within the natural video data manifold learned by pre-trained models, and the core challenge is simply to locate this optimal solution. We solve this via our core mechanism: convergent manifold alternating projection iterations that optimize the initial noise. Extensive experiments demonstrate that NeoMap significantly outperforms all existing methods across 3 standard novel view synthesis benchmarks, including the challenging Tanks-and-Temples, LLFF and DAVIS datasets, achieving state-of-the-art generation fidelity and top-tier view consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。