用音频驱动生成更自然的口型、表情和头部动作的视频人物
AI killed the video star. Audio-driven diffusion model for expressive talking head generation
- 采用3D表示的条件运动扩散变换器建模面部动态
- 在VoxCeleb2和CelebV-HQ数据集上优于现有方法
- 适合需要高真实感人脸视频生成的研究与应用
我们提出Dimitra++,一种用于音频驱动说话头生成的新框架,专注于学习唇部运动、面部表情及头部姿态变化。具体而言,我们设计了条件运动扩散变换器(cMDT)来建模面部运动序列,并采用3D表示。cMDT接收两个输入:参考面部图像(决定外观)和音频序列(驱动动作)。在VoxCeleb2和CelebV-HQ两个常用数据集上的定量、定性实验以及用户研究均表明,Dimitra++在生成具备唇动、表情与头部姿态的真实感说话头方面优于现有方法。
原文摘要 · Abstract (English)
We propose Dimitra++, a novel framework for audio-driven talking head generation, streamlined to learn lip motion, facial expression, as well as head pose motion. Specifically, we propose a conditional Motion Diffusion Transformer (cMDT) to model facial motion sequences, employing a 3D representation. The cMDT is conditioned on two inputs: a reference facial image, which determines appearance, as well as an audio sequence, which drives the motion. Quantitative and qualitative experiments, as well as a user study on two widely employed datasets, i.e., VoxCeleb2 and CelebV-HQ, suggest that Dimitra++ is able to outperform existing approaches in generating realistic talking heads imparting lip motion, facial expression, and head pose.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。