arXiv:2412.09262cs.CV2024-12被引 57

用SyncNet监督提升音频驱动模型的口型同步精度

LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision

  • 引入SyncNet监督,强制学习音视频关联性
  • 稳定版SyncNet使口型同步准确率从91%提升至94%
  • 新增时序对齐机制,显著改善生成视频的时序一致性

端到端音频条件潜空间扩散模型在音频驱动人脸动画中广泛应用,能生成逼真高分辨率说话视频。然而直接用于口型同步任务时,同步精度不佳。深入分析发现,问题源于模型倾向于学习视觉-视觉捷径,忽略关键音视频关联。为此,我们探索将SyncNet监督融入音频条件扩散模型,显式强化音视频关联学习。由于SyncNet性能直接影响同步效果,其训练收敛至关重要。我们首次开展系统性实证研究,识别影响SyncNet收敛的关键因素,并提出StableSyncNet,其架构设计保障稳定收敛。在HDTF测试集上,准确率从91%提升至94%。此外,引入新型时序表示对齐(TREPA)机制,增强生成视频时序一致性。实验表明,该方法在HDTF与VoxCeleb2数据集上,多项评估指标均优于现有先进方法。

原文摘要 · Abstract (English)

End-to-end audio-conditioned latent diffusion models (LDMs) have been widely adopted for audio-driven portrait animation, demonstrating their effectiveness in generating lifelike and high-resolution talking videos. However, direct application of audio-conditioned LDMs to lip-synchronization (lip-sync) tasks results in suboptimal lip-sync accuracy. Through an in-depth analysis, we identified the underlying cause as the "shortcut learning problem", wherein the model predominantly learns visual-visual shortcuts while neglecting the critical audio-visual correlations. To address this issue, we explored different approaches for integrating SyncNet supervision into audio-conditioned LDMs to explicitly enforce the learning of audio-visual correlations. Since the performance of SyncNet directly influences the lip-sync accuracy of the supervised model, the training of a well-converged SyncNet becomes crucial. We conducted the first comprehensive empirical studies to identify key factors affecting SyncNet convergence. Based on our analysis, we introduce StableSyncNet, with an architecture designed for stable convergence. Our StableSyncNet achieved a significant improvement in accuracy, increasing from 91% to 94% on the HDTF test set. Additionally, we introduce a novel Temporal Representation Alignment (TREPA) mechanism to enhance temporal consistency in the generated videos. Experimental results show that our method surpasses state-of-the-art lip-sync approaches across various evaluation metrics on the HDTF and VoxCeleb2 datasets.

口型同步扩散模型音视频对齐StableSyncNet

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。