arXiv:2504.18810cs.CVcs.AI2025-04被引 3

通过联合学习视觉不确定性,提升语音驱动人脸视频生成的精度与鲁棒性。

Audio-Driven Talking Face Video Generation with Joint Uncertainty Learning

  • 设计不确定性模块,同时预测图像误差与估计置信度。
  • 利用KL散度和直方图近似,使误差与不确定性分布对齐。
  • 在多种输入条件下保持高保真度与语音唇动同步性,适合数字人应用。

语音驱动的人脸视频生成是数字人技术中的重要挑战。现有方法多关注音频-唇部同步与视觉质量,却忽视了视觉不确定性的学习,导致生成结果在不同输入下表现不一、可靠性差。为此,我们提出联合不确定性学习网络(JULNet),引入与视觉误差直接相关的不确定性表示。首先,在生成图像后分别预测误差图(生成图与真实图的差异)和不确定性图(错误估计的概率)。为进一步使不确定性分布与误差分布匹配,引入基于直方图的KL散度约束。通过联合优化误差与不确定性,显著提升模型性能与鲁棒性。大量实验表明,该方法在保真度与音频-唇部同步性上均优于现有方法。

原文摘要 · Abstract (English)

Talking face video generation with arbitrary speech audio is a significant challenge within the realm of digital human technology. The previous studies have emphasized the significance of audio-lip synchronization and visual quality. Currently, limited attention has been given to the learning of visual uncertainty, which creates several issues in existing systems, including inconsistent visual quality and unreliable performance across different input conditions. To address the problem, we propose a Joint Uncertainty Learning Network (JULNet) for high-quality talking face video generation, which incorporates a representation of uncertainty that is directly related to visual error. Specifically, we first design an uncertainty module to individually predict the error map and uncertainty map after obtaining the generated image. The error map represents the difference between the generated image and the ground truth image, while the uncertainty map is used to predict the probability of incorrect estimates. Furthermore, to match the uncertainty distribution with the error distribution through a KL divergence term, we introduce a histogram technique to approximate the distributions. By jointly optimizing error and uncertainty, the performance and robustness of our model can be enhanced. Extensive experiments demonstrate that our method achieves superior high-fidelity and audio-lip synchronization in talking face video generation compared to previous methods.

人脸生成语音驱动不确定性学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。