arXiv:2502.12002cs.SDcs.CV2025-02被引 4

用可微信号处理直接生成语音,提升唇语转语音的自然度和保真度。

NaturalL2S: End-to-End High-quality Multispeaker Lip-to-Speech Synthesis with Differential Digital Signal Processing

  • 端到端设计,融合基频预测与可微信号处理生成语音
  • 避免中间频谱表示,减少合成与真实语音的域差距
  • 无需参考音色嵌入,仍能保持良好说话人相似性

近年来视觉语音识别(VSR)的进步推动了唇语转语音(L2S)的发展,预训练的VSR模型通过提供语义信息提升了合成语音的可懂度。现有级联框架虽结合伪VSR与伪文本转语音(TTS),或隐式利用转写文本,但通常依赖梅尔频谱作为中间表示,由此产生合成频谱与声码器训练所用真实频谱之间的域差距,成为影响合成质量的关键瓶颈。为解决此问题,本文提出自然唇语转语音(NaturalL2S)框架,将声学先验知识与可微语音生成组件结合。具体地,引入基频(F0)预测器以捕捉合成语音的韵律变化,预测的F0驱动可微数字信号处理(DDSP)合成器生成粗略信号,作为后续语音合成的先验信息。此外,不依赖参考说话人嵌入,仍实现良好的说话人相似性表现。客观与主观评估结果表明,NaturalL2S在合成语音质量上优于现有最先进方法。

原文摘要 · Abstract (English)

Recent advancements in visual speech recognition (VSR) have promoted progress in lip-to-speech synthesis, where pre-trained VSR models enhance the intelligibility of synthesized speech by providing valuable semantic information. The success achieved by cascade frameworks, which combine pseudo-VSR with pseudo-text-to-speech (TTS) or implicitly utilize the transcribed text, highlights the benefits of leveraging VSR models. However, these methods typically rely on mel-spectrograms as an intermediate representation, which may introduce a key bottleneck: the domain gap between synthetic mel-spectrograms, generated from inherently error-prone lip-to-speech mappings, and real mel-spectrograms used to train vocoders. This mismatch inevitably degrades synthesis quality. To bridge this gap, we propose Natural Lip-to-Speech (NaturalL2S), an end-to-end framework integrating acoustic inductive biases with differentiable speech generation components. Specifically, we introduce a fundamental frequency (F0) predictor to capture prosodic variations in synthesized speech. The predicted F0 then drives a Differentiable Digital Signal Processing (DDSP) synthesizer to generate a coarse signal which serves as prior information for subsequent speech synthesis. Additionally, instead of relying on a reference speaker embedding as an auxiliary input, our approach achieves satisfactory performance on speaker similarity without explicitly modelling speaker characteristics. Both objective and subjective evaluation results demonstrate that NaturalL2S can effectively enhance the quality of the synthesized speech when compared to state-of-the-art methods. Our demonstration page is accessible at https://yifan-liang.github.io/NaturalL2S/.

唇语转语音可微信号处理语音合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。