用单视角视频生成360度同步的真人多视角动作视频
MV-Performer: Taming Video Diffusion Model for Faithful and Synchronized Multi-view Performer Synthesis
- 基于部分点云法向图构建相机依赖条件信号,减少视图歧义
- 在三个数据集上实现当前最佳的多视角同步生成效果
- 适合需要高质量人体4D重建与跨视角动画的应用场景
近期视频生成领域的突破得益于大规模数据集和扩散技术,表明视频扩散模型可作为隐式的4D新视角合成器。然而,现有方法主要聚焦于前视图内相机轨迹调整,难以实现360度视角变化。本文聚焦于以人为中心的子领域,提出MV-Performer框架,从单视角全身捕捉生成同步的新视角视频。为实现360度合成,我们广泛利用MVHumanNet数据集并引入信息性条件信号,采用从定向局部点云渲染的相机依赖法向图,有效缓解可见与不可见观测间的歧义。为保持生成视频的同步性,提出一种多视角人本视频扩散模型,融合参考视频、局部渲染与不同视角信息。此外,针对真实环境视频,设计鲁棒推理流程,显著降低因单目深度估计不准确带来的伪影。在三个数据集上的大量实验验证了MV-Performer在效果与鲁棒性方面的先进性,为人体为中心的4D新视角合成树立了新基准。
原文摘要 · Abstract (English)
Recent breakthroughs in video generation, powered by large-scale datasets and diffusion techniques, have shown that video diffusion models can function as implicit 4D novel view synthesizers. Nevertheless, current methods primarily concentrate on redirecting camera trajectory within the front view while struggling to generate 360-degree viewpoint changes. In this paper, we focus on human-centric subdomain and present MV-Performer, an innovative framework for creating synchronized novel view videos from monocular full-body captures. To achieve a 360-degree synthesis, we extensively leverage the MVHumanNet dataset and incorporate an informative condition signal. Specifically, we use the camera-dependent normal maps rendered from oriented partial point clouds, which effectively alleviate the ambiguity between seen and unseen observations. To maintain synchronization in the generated videos, we propose a multi-view human-centric video diffusion model that fuses information from the reference video, partial rendering, and different viewpoints. Additionally, we provide a robust inference procedure for in-the-wild video cases, which greatly mitigates the artifacts induced by imperfect monocular depth estimation. Extensive experiments on three datasets demonstrate our MV-Performer's state-of-the-art effectiveness and robustness, setting a strong model for human-centric 4D novel view synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。