arXiv:2504.19165cs.CV2025-04CVPR被引 3

单张图像生成逼真会动的人脸视频,直接输出3D一致结果

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos

  • 通过单步去噪直接生成多平面图像,无需额外3D重建步骤
  • 在无多视角数据下仍能保持高质量图像与3D一致性
  • 适合虚拟主播、VR头显等需要沉浸式人脸视频的场景

我们提出一种新颖的基于扩散模型的3D-aware方法,可直接从单张身份图像和显式控制信号(如表情)生成逼真会动的人脸视频。该方法生成多平面图像(MPIs),确保几何一致性,适用于双目视频等沉浸式体验场景。与现有需独立3D重建阶段(如NeRF或3D高斯)的方法不同,本方法通过单一去噪过程直接输出最终结果,避免了渲染新视角所需的后处理。为有效从单目视频学习,我们引入一种训练机制:随机在目标或参考相机空间中重构输出的MPI,使模型同时学习清晰图像细节与底层3D信息。大量实验表明,即使缺乏显式3D重建或高质量多视角训练数据,该方法仍能实现竞争性的人物质量与新视角渲染能力。

原文摘要 · Abstract (English)

We propose a novel 3D-aware diffusion-based method for generating photorealistic talking head videos directly from a single identity image and explicit control signals (e.g., expressions). Our method generates Multiplane Images (MPIs) that ensure geometric consistency, making them ideal for immersive viewing experiences like binocular videos for VR headsets. Unlike existing methods that often require a separate stage or joint optimization to reconstruct a 3D representation (such as NeRF or 3D Gaussians), our approach directly generates the final output through a single denoising process, eliminating the need for post-processing steps to render novel views efficiently. To effectively learn from monocular videos, we introduce a training mechanism that reconstructs the output MPI randomly in either the target or the reference camera space. This approach enables the model to simultaneously learn sharp image details and underlying 3D information. Extensive experiments demonstrate the effectiveness of our method, which achieves competitive avatar quality and novel-view rendering capabilities, even without explicit 3D reconstruction or high-quality multi-view training data.

视频生成扩散模型3D感知人脸动画

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。