arXiv:2510.12785cs.CVcs.AI2025-10SIGGRAPH被引 17

仅用一张照片生成可动画化的360度数字人视频

MVP4D: Multi-View Portrait Video Diffusion for Animatable 4D Avatars

  • 基于预训练视频扩散模型,从单张参考图生成多视角动态视频
  • 支持360度视角生成,画面真实度与时间一致性显著提升
  • 适合游戏、虚拟演出等需要高逼真数字人的场景

数字人化身旨在虚拟环境中模拟人类的动态外观,广泛应用于游戏、影视和虚拟现实等领域。传统制作高质量数字人耗时耗力,需大型摄像机阵列及专业3D艺术家手动建模。近年来,随着图像与视频生成模型的发展,已有方法可从单张随意拍摄的参考图像自动生成逼真的动画化身。但这些方法缺乏多视角信息或显式3D表示,导致视角偏离参考图时画质与真实感下降。本文提出MVP4D,基于先进预训练视频扩散模型,从单张参考图生成覆盖最多360度视角的动画视频,可同时输出数百帧。我们进一步将模型输出蒸馏为可实时渲染的4D化身。相比此前方法,本方案在真实感、时间一致性和3D一致性上均有显著提升。

原文摘要 · Abstract (English)

Digital human avatars aim to simulate the dynamic appearance of humans in virtual environments, enabling immersive experiences across gaming, film, virtual reality, and more. However, the conventional process for creating and animating photorealistic human avatars is expensive and time-consuming, requiring large camera capture rigs and significant manual effort from professional 3D artists. With the advent of capable image and video generation models, recent methods enable automatic rendering of realistic animated avatars from a single casually captured reference image of a target subject. While these techniques significantly lower barriers to avatar creation and offer compelling realism, they lack constraints provided by multi-view information or an explicit 3D representation. So, image quality and realism degrade when rendered from viewpoints that deviate strongly from the reference image. Here, we build a video model that generates animatable multi-view videos of digital humans based on a single reference image and target expressions. Our model, MVP4D, is based on a state-of-the-art pre-trained video diffusion model and generates hundreds of frames simultaneously from viewpoints varying by up to 360 degrees around a target subject. We show how to distill the outputs of this model into a 4D avatar that can be rendered in real-time. Our approach significantly improves the realism, temporal consistency, and 3D consistency of generated avatars compared to previous methods.

数字人视频生成扩散模型360度渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。