用音频精准控制3D人脸口型,生成更自然的说话视频。
JoyGen: Audio-Driven 3D Depth-Aware Talking-Face Video Editing
- 分两阶段生成:先预测表情参数,再结合深度图精调口型动作。
- 在130小时中文数据集上训练,口型与音频同步率显著提升。
- 适合需要高保真语音驱动人脸生成的研究者和开发者。
说话人脸视频生成研究已取得显著进展,但基于输入音频精确编辑唇部动作并保持高视觉质量仍是挑战。本文提出JoyGen,一种两阶段框架,包含音频驱动唇部运动生成与视觉外观合成。第一阶段通过3D重建模型和audio2motion模型分别预测身份与表情系数。随后,结合音频特征与面部深度图,提供全面监督以实现面部生成中的精确唇音同步。此外,我们构建了一个包含130小时高质量视频的中文说话人脸数据集。JoyGen在开源HDTF数据集及自建数据集上进行训练。实验结果表明,该方法在唇音同步性和视觉质量方面均表现优异。
原文摘要 · Abstract (English)
Significant progress has been made in talking-face video generation research; however, precise lip-audio synchronization and high visual quality remain challenging in editing lip shapes based on input audio. This paper introduces JoyGen, a novel two-stage framework for talking-face generation, comprising audio-driven lip motion generation and visual appearance synthesis. In the first stage, a 3D reconstruction model and an audio2motion model predict identity and expression coefficients respectively. Next, by integrating audio features with a facial depth map, we provide comprehensive supervision for precise lip-audio synchronization in facial generation. Additionally, we constructed a Chinese talking-face dataset containing 130 hours of high-quality video. JoyGen is trained on the open-source HDTF dataset and our curated dataset. Experimental results demonstrate superior lip-audio synchronization and visual quality achieved by our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。