arXiv:2603.19613cs.CV2026-03被引 1

用视频生成模型做新视角合成,提升单视图下的真实感和几何一致性。

OrbitNVS: Harnessing Video Diffusion Priors for Novel View Synthesis

论文配图:OrbitNVS: Harnessing Video Diffusion Priors for Novel View Synthesis
图 1 · 摘自论文原文
  • 将新视角合成转化为轨道视频生成任务,利用预训练视频模型先验。
  • 在单视图下达到2.9dB和2.4dB的PSNR提升,显著优于以往方法。
  • 适合需要高质量3D视图生成的研究者或工业应用开发者。

新视角合成(NVS)旨在仅凭有限已知视角生成目标3D物体的未见视角。现有方法在单视图输入下难以合成合理未观测区域,且常面临几何与外观不一致的问题。为此,本文提出OrbitNVS,将NVS重构为轨道视频生成任务。通过定制化模型设计与训练策略,我们适配预训练视频生成模型以利用其丰富的视觉先验,实现高质量视图合成。具体地,引入相机适配器以实现精确相机控制;设计法向图生成分支,并利用法向图特征通过注意力机制引导目标视图合成,提升几何一致性;此外,采用像素空间监督缓解潜在空间压缩导致的模糊问题。大量实验表明,OrbitNVS在GSO和OmniObject3D基准上显著优于先前方法,尤其在挑战性的单视图设置下(例如,PSNR提升+2.9 dB和+2.4 dB)。

原文摘要 · Abstract (English)

Novel View Synthesis (NVS) aims to generate unseen views of a 3D object given a limited number of known views. Existing methods often struggle to synthesize plausible views for unobserved regions, particularly under single-view input, and still face challenges in maintaining geometry- and appearance-consistency. To address these issues, we propose OrbitNVS, which reformulates NVS as an orbit video generation task. Through tailored model design and training strategies, we adapt a pre-trained video generation model to the NVS task, leveraging its rich visual priors to achieve high-quality view synthesis. Specifically, we incorporate camera adapters into the video model to enable accurate camera control. To enhance two key properties of 3D objects, geometry and appearance, we design a normal map generation branch and use normal map features to guide the synthesis of the target views via attention mechanism, thereby improving geometric consistency. Moreover, we apply a pixel-space supervision to alleviate blurry appearance caused by spatial compression in the latent space. Extensive experiments show that OrbitNVS significantly outperforms previous methods on the GSO and OmniObject3D benchmarks, especially in the challenging single-view setting (\eg, +2.9 dB and +2.4 dB PSNR).

新视角合成视频生成扩散模型3D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。