用音频生成更自然的虚拟人脸,支持口型、表情和头部动作同步。
Dimitra: Audio-driven Diffusion model for Expressive Talking Head Generation
- 基于3D表示的条件运动扩散模型,仅需音频和参考图像输入。
- 在VoxCeleb2和HDTF数据集上优于现有方法,提升口型与表情真实度。
- 音频中音素影响口型,文本信息影响表情和头部姿态,机制清晰。
我们提出Dimitra,一种新型音频驱动的人脸生成框架,专注于学习唇部动作、面部表情及头部姿态运动。具体而言,通过3D表示建模面部运动序列,训练一个条件运动扩散变换器(cMDT)。cMDT仅以音频序列和参考面部图像为输入,直接从音频中提取额外特征,从而提升生成视频的质量与真实感。其中,音素序列增强口型真实性,文本转录内容提升表情与头部姿态表现力。在VoxCeleb2和HDTF两个常用数据集上的定量与定性实验表明,Dimitra在生成具备口型、表情和头部姿态的逼真人脸方面优于现有方法。
原文摘要 · Abstract (English)
We propose Dimitra, a novel framework for audio-driven talking head generation, streamlined to learn lip motion, facial expression, as well as head pose motion. Specifically, we train a conditional Motion Diffusion Transformer (cMDT) by modeling facial motion sequences with 3D representation. We condition the cMDT with only two input signals, an audio-sequence, as well as a reference facial image. By extracting additional features directly from audio, Dimitra is able to increase quality and realism of generated videos. In particular, phoneme sequences contribute to the realism of lip motion, whereas text transcript to facial expression and head pose realism. Quantitative and qualitative experiments on two widely employed datasets, VoxCeleb2 and HDTF, showcase that Dimitra is able to outperform existing approaches for generating realistic talking heads imparting lip motion, facial expression, and head pose.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。