用一张自拍视频实时生成细节丰富的头部虚拟形象。
SelfieAvatar: Real-time Head Avatar reenactment from a Selfie Video
- 结合3DMM与StyleGAN,通过混合损失函数重建高细节头像。
- 在自重演和跨重演任务中均超越现有方法,纹理更精细。
- 仅需一张自拍视频训练,适合个性化虚拟形象应用。
头部虚拟形象重演旨在从单目视频中创建可驱动的个性化虚拟形象,是社交信号理解、游戏、人机交互及计算机视觉的基础。近期基于3D可变形模型(3DMM)的面部重建方法已实现高保真度人脸估计。然而,这些方法难以实时捕捉完整头部(包括非面部区域和背景细节),而基于生成对抗网络(GAN)的方法虽能生成高质量重演效果,却难以还原皱纹、发丝等细粒度特征。此外,现有方法通常依赖大量训练数据,极少关注仅用简单自拍视频实现虚拟形象重演。为此,本文提出一种基于自拍视频的详细头部虚拟形象重演方法,结合3DMM与基于StyleGAN的生成器。设计了融合前景重建与虚拟图像生成的混合损失函数,在对抗训练中恢复高频细节。在自重演与跨重演任务上的定性与定量评估表明,该方法在头部虚拟形象重建上表现更优,纹理更丰富、更细腻。
原文摘要 · Abstract (English)
Head avatar reenactment focuses on creating animatable personal avatars from monocular videos, serving as a foundational element for applications like social signal understanding, gaming, human-machine interaction, and computer vision. Recent advances in 3D Morphable Model (3DMM)-based facial reconstruction methods have achieved remarkable high-fidelity face estimation. However, on the one hand, they struggle to capture the entire head, including non-facial regions and background details in real time, which is an essential aspect for producing realistic, high-fidelity head avatars. On the other hand, recent approaches leveraging generative adversarial networks (GANs) for head avatar generation from videos can achieve high-quality reenactments but encounter limitations in reproducing fine-grained head details, such as wrinkles and hair textures. In addition, existing methods generally rely on a large amount of training data, and rarely focus on using only a simple selfie video to achieve avatar reenactment. To address these challenges, this study introduces a method for detailed head avatar reenactment using a selfie video. The approach combines 3DMMs with a StyleGAN-based generator. A detailed reconstruction model is proposed, incorporating mixed loss functions for foreground reconstruction and avatar image generation during adversarial training to recover high-frequency details. Qualitative and quantitative evaluations on self-reenactment and cross-reenactment tasks demonstrate that the proposed method achieves superior head avatar reconstruction with rich and intricate textures compared to existing approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。