用分层扩散模型实现音视频人脸同步生成,效果逼真且适配多身份。
DreamHead: Learning Spatial-Temporal Correspondence via Hierarchical Diffusion for Audio-driven Talking Head Synthesis
- 分两阶段:先由音频预测连续面部关键点,再由关键点生成连贯人脸视频。
- 在多个身份上生成高质量语音驱动人脸视频,保持时空一致性。
- 无需额外训练即可适配不同人物,适合跨身份语音动画应用。
语音驱动的人脸视频合成旨在从输入音频生成逼真的人物肖像视频。扩散模型因其出色的生成质量与强泛化能力,被用于该任务。然而,如何在扩散模型中建立音频时序特征与对应面部空间表情之间的鲁棒映射,仍是核心挑战。为此,我们提出 DreamHead,一种分层扩散框架,可在不损害模型内在质量与适应性的情况下学习时空对应关系。DreamHead 将密集面部关键点作为中间信号,首先设计了一个音频到关键点的扩散过程,以生成时间平滑且准确的关键点序列;随后,进一步提出一个关键点到图像的扩散过程,通过建模关键点与外观之间的空间对应关系,生成空间一致的面部视频。大量实验表明,所提出的 DreamHead 能有效学习时空一致性,并为多种身份生成高保真语音驱动的人脸视频。
原文摘要 · Abstract (English)
Audio-driven talking head synthesis strives to generate lifelike video portraits from provided audio. The diffusion model, recognized for its superior quality and robust generalization, has been explored for this task. However, establishing a robust correspondence between temporal audio cues and corresponding spatial facial expressions with diffusion models remains a significant challenge in talking head generation. To bridge this gap, we present DreamHead, a hierarchical diffusion framework that learns spatial-temporal correspondences in talking head synthesis without compromising the model's intrinsic quality and adaptability.~DreamHead learns to predict dense facial landmarks from audios as intermediate signals to model the spatial and temporal correspondences.~Specifically, a first hierarchy of audio-to-landmark diffusion is first designed to predict temporally smooth and accurate landmark sequences given audio sequence signals. Then, a second hierarchy of landmark-to-image diffusion is further proposed to produce spatially consistent facial portrait videos, by modeling spatial correspondences between the dense facial landmark and appearance. Extensive experiments show that proposed DreamHead can effectively learn spatial-temporal consistency with the designed hierarchical diffusion and produce high-fidelity audio-driven talking head videos for multiple identities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。