arXiv:2410.06734cs.CV2024-10NeurIPS被引 25

15分钟生成个性化高保真3D说话脸,效率提升47倍

MimicTalk: Mimicking a personalized and expressive 3D talking face in minutes

  • 用通用3D模型+混合静态动态适配,快速定制个人形象
  • 参考视频风格迁移,生成自然表情与口型动作
  • 适合影视动画、虚拟主播等需快速生成个性化角色的场景

说话人脸生成(TFG)旨在为特定人物生成逼真的说话视频。个性化TFG强调合成结果在外观和说话风格上的感知一致性。现有方法通常为每个身份训练独立的神经辐射场(NeRF),但存在效率低、泛化差的问题。为此,我们提出MimicTalk,首次利用基于NeRF的通用无身份模型知识,提升个性化TFG的效率与鲁棒性。具体包括:(1) 构建无身份的3D TFG基础模型,并可快速适配至特定身份;(2) 提出静态-动态混合适配流程,学习个性化外观与面部动态特征;(3) 设计上下文风格化音频到运动模型,通过隐式风格表示实现参考视频说话风格的无损模仿。对未见身份的适配仅需15分钟,较以往方法快47倍。实验表明,MimicTalk在视频质量、效率与表现力上均优于基线。源码与样例视频见https://mimictalk.github.io。

原文摘要 · Abstract (English)

Talking face generation (TFG) aims to animate a target identity's face to create realistic talking videos. Personalized TFG is a variant that emphasizes the perceptual identity similarity of the synthesized result (from the perspective of appearance and talking style). While previous works typically solve this problem by learning an individual neural radiance field (NeRF) for each identity to implicitly store its static and dynamic information, we find it inefficient and non-generalized due to the per-identity-per-training framework and the limited training data. To this end, we propose MimicTalk, the first attempt that exploits the rich knowledge from a NeRF-based person-agnostic generic model for improving the efficiency and robustness of personalized TFG. To be specific, (1) we first come up with a person-agnostic 3D TFG model as the base model and propose to adapt it into a specific identity; (2) we propose a static-dynamic-hybrid adaptation pipeline to help the model learn the personalized static appearance and facial dynamic features; (3) To generate the facial motion of the personalized talking style, we propose an in-context stylized audio-to-motion model that mimics the implicit talking style provided in the reference video without information loss by an explicit style representation. The adaptation process to an unseen identity can be performed in 15 minutes, which is 47 times faster than previous person-dependent methods. Experiments show that our MimicTalk surpasses previous baselines regarding video quality, efficiency, and expressiveness. Source code and video samples are available at https://mimictalk.github.io .

3D说话脸个性化生成高效适配风格迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。