arXiv:2411.11258cs.SDeess.AS2024-11中稿 · NCMMSC2024

用音调与频谱先验信息提升语音合成质量与速度

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram

  • 基于声源-滤波理论,用神经滤波器将激励频谱转为语音频谱
  • 在多个数据集上优于或接近HiFi-GAN等基线模型,生成速度快
  • 激励结构加速收敛,适合追求高效高质语音合成的研究者

本文提出ESTVocoder,一种基于声源-滤波理论的新型神经声码器。该模型利用以ConvNeXt v2块为骨干的神经滤波器,将激励的幅值和相位谱转换为对应的语音幅值和相位谱,再通过逆短时傅里叶变换(ISTFT)重建语音波形。激励基于基频(F0)构建:有声段包含完整谐波信息,无声段则用噪声表示。激励为滤波器提供幅度和相位模式的先验知识,降低建模难度。为保证合成语音保真度,采用多尺度、多分辨率判别器进行对抗训练。分析-合成及文语转换实验表明,本方法在语音质量上优于或媲美HiFi-GAN、SiFi-GAN和Vocos等基线模型,且具有合理的模型复杂度和生成速度。额外分析显示,引入的激励有效加速了模型收敛,得益于激励中蕴含的语音频谱先验信息。

原文摘要 · Abstract (English)

This paper proposes ESTVocoder, a novel excitation-spectral-transformed neural vocoder within the framework of source-filter theory. The ESTVocoder transforms the amplitude and phase spectra of the excitation into the corresponding speech amplitude and phase spectra using a neural filter whose backbone is ConvNeXt v2 blocks. Finally, the speech waveform is reconstructed through the inverse short-time Fourier transform (ISTFT). The excitation is constructed based on the F0: for voiced segments, it contains full harmonic information, while for unvoiced segments, it is represented by noise. The excitation provides the filter with prior knowledge of the amplitude and phase patterns, expecting to reduce the modeling difficulty compared to conventional neural vocoders. To ensure the fidelity of the synthesized speech, an adversarial training strategy is applied to ESTVocoder with multi-scale and multi-resolution discriminators. Analysis-synthesis and text-to-speech experiments both confirm that our proposed ESTVocoder outperforms or is comparable to other baseline neural vocoders, e.g., HiFi-GAN, SiFi-GAN, and Vocos, in terms of synthesized speech quality, with a reasonable model complexity and generation speed. Additional analysis experiments also demonstrate that the introduced excitation effectively accelerates the model's convergence process, thanks to the speech spectral prior information contained in the excitation.

语音合成神经声码器声源滤波

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。