arXiv:2503.13371cs.LG2025-03中稿 · WACV 2025被引 2

用时序姿态与音频特征提升扩散模型的口型同步性

SyncDiff: Diffusion-based Talking Head Synthesis with Bottlenecked Temporal Visual Prior for Improved Synchronization

  • 引入带信息瓶颈的时序姿态帧和音频特征作为扩散模型条件输入
  • 在LRS2/LRS3上同步得分分别提升27.7%/62.3%,保持高画质
  • 适合需要高同步性与高图像质量的语音驱动人脸生成场景

说话头合成(即语音到口型合成)旨在重建与给定音频对齐的面部动作。合成视频主要从口型-语音同步性和图像保真度两方面评估。近期研究表明,基于GAN和扩散模型的方法在该任务上达到当前最佳性能,其中扩散模型虽具备更优的图像保真度,但同步性低于基于GAN的方法。为此,本文提出SyncDiff,一种简单而有效的方法:通过引入带有信息瓶颈的时序姿态帧以及从AVHuBERT提取的面部相关音频特征,作为扩散过程的条件输入。我们在两个标准说话头数据集LRS2和LRS3上进行评估,并与现有SOTA模型直接对比。实验表明,SyncDiff在LRS2/LRS3上的同步分数相较于先前扩散模型分别提升了27.7%/62.3%,同时保留了其高保真特性。

原文摘要 · Abstract (English)

Talking head synthesis, also known as speech-to-lip synthesis, reconstructs the facial motions that align with the given audio tracks. The synthesized videos are evaluated on mainly two aspects, lip-speech synchronization and image fidelity. Recent studies demonstrate that GAN-based and diffusion-based models achieve state-of-the-art (SOTA) performance on this task, with diffusion-based models achieving superior image fidelity but experiencing lower synchronization compared to their GAN-based counterparts. To this end, we propose SyncDiff, a simple yet effective approach to improve diffusion-based models using a temporal pose frame with information bottleneck and facial-informative audio features extracted from AVHuBERT, as conditioning input into the diffusion process. We evaluate SyncDiff on two canonical talking head datasets, LRS2 and LRS3 for direct comparison with other SOTA models. Experiments on LRS2/LRS3 datasets show that SyncDiff achieves a synchronization score 27.7%/62.3% relatively higher than previous diffusion-based methods, while preserving their high-fidelity characteristics.

说话头合成扩散模型口型同步音视频对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。