arXiv:2511.22293cs.SDcs.LG2025-11

改进语音合成中的相位对齐,提升生成音频质量与速度

GLA-Grad++: An Improved Griffin-Lim Guided Diffusion Model for Speech Synthesis

  • 在反向扩散过程中仅用一次格里芬-林算法修正相位
  • 在域外数据上音质显著优于基线模型
  • 适合需要快速高保真语音合成的场景

扩散模型在语音合成中展现出强大生成能力,显著提升音频质量与稳定性。然而,当条件输入(梅尔频谱图)偏离训练分布时,其表现受限。近期提出的GLA-Grad模型通过在波形生成反向过程中引入格里芬-林算法(GLA),增强了生成信号与条件频谱的一致性。本文进一步改进该方法:仅在生成过程开始时进行一次GLA校正,大幅加速生成速度。实验表明,该方法在各类测试场景下均优于基线模型,尤其在域外数据上表现更优。

原文摘要 · Abstract (English)

Recent advances in diffusion models have positioned them as powerful generative frameworks for speech synthesis, demonstrating substantial improvements in audio quality and stability. Nevertheless, their effectiveness in vocoders conditioned on mel spectrograms remains constrained, particularly when the conditioning diverges from the training distribution. The recently proposed GLA-Grad model introduced a phase-aware extension to the WaveGrad vocoder that integrated the Griffin-Lim algorithm (GLA) into the reverse process to reduce inconsistencies between generated signals and conditioning mel spectrogram. In this paper, we further improve GLA-Grad through an innovative choice in how to apply the correction. Particularly, we compute the correction term only once, with a single application of GLA, to accelerate the generation process. Experimental results demonstrate that our method consistently outperforms the baseline models, particularly in out-of-domain scenarios.

语音合成扩散模型相位校正

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。