arXiv:2505.11391eess.AScs.SD2025-05被引 4

用扩散模型从口型视频生成自然语音,音质和说话人相似度更优。

LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models

  • 基于条件扩散模型,结合视觉特征与说话人嵌入进行语音生成。
  • 在LRS3数据集上语音感知质量与说话人相似度超越现有方法。
  • 适合需要高质量口型转语音的应用场景,如影视配音、无障碍沟通。

我们提出LipDiffuser,一种用于口型转语音的条件扩散模型,可直接从无声视频生成自然且可懂的语音。该方法采用保持幅度的消融扩散模型(MP-ADM)作为去噪器,并通过保持幅度的特征线性调制(MP-FiLM)结合视觉特征与说话人嵌入以实现有效条件控制。生成的梅尔频谱图由神经声码器还原为语音波形。在LRS3数据集上的评估显示,LipDiffuser在语音感知质量与说话人相似度方面优于现有基线,同时在下游自动语音识别任务中仍具竞争力。这些结果也得到了正式听觉实验的支持。

原文摘要 · Abstract (English)

We present LipDiffuser, a conditional diffusion model for lip-to-speech generation synthesizing natural and intelligible speech directly from silent video recordings. Our approach leverages the magnitude-preserving ablated diffusion model (MP-ADM) architecture as a denoiser model. To effectively condition the model, we incorporate visual features using magnitude-preserving feature-wise linear modulation (MP-FiLM) alongside speaker embeddings. A neural vocoder then reconstructs the speech waveform from the generated mel-spectrograms. Evaluations on LRS3 demonstrate that LipDiffuser outperforms existing lip-to-speech baselines in perceptual speech quality and speaker similarity, while remaining competitive in downstream automatic speech recognition. These findings are also supported by a formal listening experiment.

口型转语音扩散模型语音合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。