arXiv:2507.01611eess.AScs.SD2025-07被引 1

将语音共振特性建模为ARMA函数,实现高效高质语音合成与灵活修改。

QHARMA-GAN: Quasi-Harmonic Neural Vocoder based on Autoregressive Moving Average Model

  • 用ARMA模型建模语音共振特征,结合神经网络提升可解释性。
  • 生成速度更快、网络更小,且支持高质量的变调和变速处理。
  • 适合需要语音灵活调控的场景,如语音编辑与个性化合成。

声码器通过将语音信号编码为声学特征并重建语音信号,已有数十年研究历史。近年来深度学习推动了神经声码器的发展,但现有端到端方法存在黑箱问题,难以体现语音生成机制与内在结构,导致声源激励与共振特性建模模糊,且难以灵活修改语音。此外,序列式波形生成常需复杂网络,耗时较长。受准谐波模型(QHM)启发,该工作提出一种融合神经网络与QHM合成过程的新框架。语音信号被编码为自回归移动平均(ARMA)函数以建模共振特性,从而在任意频率上精确估计准谐波的振幅与相位。随后可高质量重合成语音,并灵活进行变调与变速操作,同时显著降低计算时间与模型规模。实验表明,该方法融合了QHM、ARMA模型与神经网络的优势,在生成速度、合成质量与修改灵活性方面均优于现有方法。

原文摘要 · Abstract (English)

Vocoders, encoding speech signals into acoustic features and allowing for speech signal reconstruction from them, have been studied for decades. Recently, the rise of deep learning has particularly driven the development of neural vocoders to generate high-quality speech signals. On the other hand, the existing end-to-end neural vocoders suffer from a black-box nature that blinds the speech production mechanism and the intrinsic structure of speech, resulting in the ambiguity of separately modeling source excitation and resonance characteristics and the loss of flexibly synthesizing or modifying speech with high quality. Moreover, their sequence-wise waveform generation usually requires complicated networks, leading to substantial time consumption. In this work, inspired by the quasi-harmonic model (QHM) that represents speech as sparse components, we combine the neural network and QHM synthesis process to propose a novel framework for the neural vocoder. Accordingly, speech signals can be encoded into autoregressive moving average (ARMA) functions to model the resonance characteristics, yielding accurate estimates of the amplitudes and phases of quasi-harmonics at any frequency. Subsequently, the speech can be resynthesized and arbitrarily modified in terms of pitch shifting and time stretching with high quality, whereas the time consumption and network size decrease. The experiments indicate that the proposed method leverages the strengths of QHM, the ARMA model, and neural networks, leading to the outperformance of our methods over other methods in terms of generation speed, synthesis quality, and modification flexibility.

语音合成神经声码器语音编辑ARMA模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。