用视频扩散模型实现单目输入下的大视角变化视图合成
UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

- 结合显式3D引导与扩散模型,提升视角生成精度
- 在极端相机运动下仍保持几何一致性和视觉保真度
- 适合需要高质量动态3D重建的影视创作场景
社交媒体上大量随意拍摄的单目视频和图像为沉浸式内容创作提供了丰富素材。然而,在输入覆盖极有限的情况下,生成逼真且几何一致的新视角并实现精确相机控制仍具挑战。基于重建的方法如NeRF和3DGS在稀疏输入下性能严重退化,无法有效处理遮挡;生成方法虽降低数据需求,但在大基线视图合成中仍因几何引导不准确而表现不佳。为此,我们提出UniWorld-View,一个统一框架,支持从单目输入进行可控的大基线新视角合成。该方法通过感知遮挡的点云渲染策略获取显式3D引导,解决可见性歧义,并为扩散模型提供精准先验。结合强大的视频扩散主干网络,UniWorld-View在极端相机运动和宽基线变化下仍能生成高保真新视角,还可输出多视角视频以支持下游动态3DGS重建。在WorldScore和零样本新视角合成基准上的实验表明,该方法在可控性、几何一致性和视觉保真度方面均表现出色。
原文摘要 · Abstract (English)
The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experiences. However, producing photorealistic and geometrically consistent views with precise camera control remains challenging when input coverage is extremely limited. Reconstruction-based approaches such as NeRF and 3D Gaussian Splatting (3DGS) deteriorate severely under sparse inputs and fail to explicitly handle occlusions. Generative methods ease data requirements but still struggle with large-baseline view synthesis due to inaccurate or implicit geometric guidance. To overcome these limitations, we introduce UniWorld-View, a unified framework for controllable large-baseline novel view synthesis from monocular inputs. UniWorld-View integrates explicit 3D guidance with generative diffusion modeling to enable precise camera control and geometrically consistent view generation. The geometric guidance is obtained through an occlusion-aware point cloud rendering strategy that resolves visibility ambiguities and provides accurate priors for diffusion-based synthesis. By coupling this rendering strategy with powerful video diffusion backbones, UniWorld-View achieves high-fidelity novel view generation even under extreme camera motions and wide-baseline changes, and can further provide multi-view videos for downstream dynamic 3DGS reconstruction. Experiments on the WorldScore benchmark and zero-shot NVS benchmarks demonstrate the effectiveness of UniWorld-View in controllability, geometric consistency, and visual fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。