用一张图生成能随声音自然动的3D人脸,细节逼真、实时渲染。
VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single Image
- 基于2D运动隐变量,构建条件3D头像模型实现表情还原。
- 单图驱动生成512x512自由视角视频,最高75帧/秒,支持实时交互。
- 适用于虚拟主播、数字人应用,对动画师和开发者友好。
我们提出 VASA-3D,一种由音频驱动、单图像输入的3D头像生成方法。该研究解决两大挑战:捕捉真实人脸中的细微表情变化,以及从单张肖像图重建精细3D头像。为精准建模表情细节,VASA-3D采用来自 VASA-1 的运动隐变量,该方法在2D动态人脸中表现出极高的真实感与生动性。本工作关键在于将此运动隐变量映射至3D空间,通过设计一个以运动隐变量为条件的3D头像模型实现。利用从输入图像合成的多帧参考头像视频,通过优化框架完成模型定制。该优化融合多种训练损失,有效应对生成数据中伪影及姿态覆盖不足的问题。实验表明,VASA-3D可生成现有技术无法实现的真实3D说话头像,支持512x512自由视角视频在线生成,最高达75 FPS,显著提升与逼真3D头像的沉浸式交互体验。
原文摘要 · Abstract (English)
We propose VASA-3D, an audio-driven, single-shot 3D head avatar generator. This research tackles two major challenges: capturing the subtle expression details present in real human faces, and reconstructing an intricate 3D head avatar from a single portrait image. To accurately model expression details, VASA-3D leverages the motion latent of VASA-1, a method that yields exceptional realism and vividness in 2D talking heads. A critical element of our work is translating this motion latent to 3D, which is accomplished by devising a 3D head model that is conditioned on the motion latent. Customization of this model to a single image is achieved through an optimization framework that employs numerous video frames of the reference head synthesized from the input image. The optimization takes various training losses robust to artifacts and limited pose coverage in the generated training data. Our experiment shows that VASA-3D produces realistic 3D talking heads that cannot be achieved by prior art, and it supports the online generation of 512x512 free-viewpoint videos at up to 75 FPS, facilitating more immersive engagements with lifelike 3D avatars.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。