arXiv:2412.19860cs.CV2024-12被引 8

UniAvatar实现高保真语音驱动人脸动画的全方位控制

UniAvatar: Taming Lifelike Audio-Driven Talking Head Generation with Comprehensive Motion and Lighting Control

  • 用FLAME模型将3D动作信息融合到单图,实现像素级精细控制
  • 支持大范围头部动作与复杂光照变化,生成效果更自然逼真
  • 适用于影视特效、虚拟主播等需要精细动画控制的场景

近期,基于音频驱动的人脸动画成为热门任务。生成逼真说话头像视频需灵活自然的面部与头部运动、摄像机位移以及真实的光影效果。现有方法难以全面控制这些多维度要素。本文提出UniAvatar,通过使用FLAME模型将所有运动信息渲染至单一图像,保留3D运动细节的同时实现像素级精细控制。该方法还支持全局光照的综合调控,设计独立模块分别管理3D运动与光照,支持单独或联合控制。大量实验表明,本方法在运动与光照控制方面均优于现有技术。此外,为提升数据集的多样性,我们收集并计划公开两个新数据集:DH-FaceDrasMvVid-100(包含100个显著头部动作的语音视频)和DH-FaceReliVid-200(涵盖200个不同光照场景的视频),以促进相关研究发展。

原文摘要 · Abstract (English)

Recently, animating portrait images using audio input is a popular task. Creating lifelike talking head videos requires flexible and natural movements, including facial and head dynamics, camera motion, realistic light and shadow effects. Existing methods struggle to offer comprehensive, multifaceted control over these aspects. In this work, we introduce UniAvatar, a designed method that provides extensive control over a wide range of motion and illumination conditions. Specifically, we use the FLAME model to render all motion information onto a single image, maintaining the integrity of 3D motion details while enabling fine-grained, pixel-level control. Beyond motion, this approach also allows for comprehensive global illumination control. We design independent modules to manage both 3D motion and illumination, permitting separate and combined control. Extensive experiments demonstrate that our method outperforms others in both broad-range motion control and lighting control. Additionally, to enhance the diversity of motion and environmental contexts in current datasets, we collect and plan to publicly release two datasets, DH-FaceDrasMvVid-100 and DH-FaceReliVid-200, which capture significant head movements during speech and various lighting scenarios.

人脸动画语音驱动光照控制3D建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。