arXiv:2506.08004cs.CVcs.AI2025-06NeurIPS被引 11

无需训练即可高保真生成动态视角,靠重构噪声实现。

Dynamic View Synthesis as an Inverse Problem

论文配图:Dynamic View Synthesis as an Inverse Problem
图 1 · 摘自论文原文
  • 重构噪声表示实现精准潜空间对齐。
  • 零终端信噪比导致反演困难,新方法有效解决。
  • 适合无训练场景下视频生成与视觉重建研究者。

本文将单目视频的动态视图合成问题视为无训练条件下的逆问题。通过重新设计预训练视频扩散模型的噪声初始化阶段,实现了无需权重更新或附加模块的高保真动态视图合成。我们识别出因零终端信噪比(SNR)调度导致的确定性反演障碍,并提出一种新型噪声表示——K阶递归噪声表示(K-order Recursive Noise Representation),推导出其闭式表达式,从而实现VAE编码潜变量与DDIM反演潜变量之间的精确高效对齐。为合成相机运动带来的新可见区域,引入随机潜变量调制(Stochastic Latent Modulation),在潜空间中进行可见性感知采样以填补遮挡区域。大量实验表明,通过在噪声初始化阶段的结构化潜空间操作,可有效完成动态视图合成。

原文摘要 · Abstract (English)

In this work, we address dynamic view synthesis from monocular videos as an inverse problem in a training-free setting. By redesigning the noise initialization phase of a pre-trained video diffusion model, we enable high-fidelity dynamic view synthesis without any weight updates or auxiliary modules. We begin by identifying a fundamental obstacle to deterministic inversion arising from zero-terminal signal-to-noise ratio (SNR) schedules and resolve it by introducing a novel noise representation, termed K-order Recursive Noise Representation. We derive a closed form expression for this representation, enabling precise and efficient alignment between the VAE-encoded and the DDIM inverted latents. To synthesize newly visible regions resulting from camera motion, we introduce Stochastic Latent Modulation, which performs visibility aware sampling over the latent space to complete occluded regions. Comprehensive experiments demonstrate that dynamic view synthesis can be effectively performed through structured latent manipulation in the noise initialization phase.

视图合成扩散模型无训练潜空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。