arXiv:2410.04704cs.SDcs.CL2024-10

用新模型更准地还原语音中的声带和声道参数,尤其适合鼻音等复杂发音。

Modeling and Estimation of Vocal Tract and Glottal Source Parameters Using ARMAX-LF Model

  • 结合自回归滑动平均与声门源模型,实现声道的极零滤波建模
  • 用深度网络直接映射语音波形到参数,误差更低且无需迭代
  • 适用于元音和鼻化辅音,实测速度更快、精度更高

从原始语音中建模与估计元音的声道与声门源参数,传统方法多采用自回归外生输入(ARX)模型与Liljencrants-Fant(LF)模型,通过迭代策略进行估计。然而,声道滤波器的全极点自回归模型无法表示反共振峰(零点),导致在鼻音、擦音和塞音等语音类型中估计误差增大。本文提出自回归滑动平均外生输入与LF模型结合的ARMAX-LF模型,扩展了对多种语音的建模能力,包括元音和鼻化辅音。其中,LF模型以时域参数化方式表示声门源导数,而ARMAX模型将声道建模为含额外外生LF激励的极零滤波器。为减少参数估计误差,首先利用深度神经网络(DNN)强大的非线性拟合能力,建立从提取的声门源导数或语音波形到对应LF参数的映射关系。由此可直接估计声门源与声道参数,避免分析-合成策略中的迭代过程。在基于线性源-滤波模型、物理模型生成的合成语音及真实语音信号上的实验表明,所提ARMAX-LF模型结合DNN估计方法,在元音与鼻化音的参数估计上均具有更低误差与更短耗时。

原文摘要 · Abstract (English)

Modeling and estimation of the vocal tract and glottal source parameters of vowels from raw speech can be typically done by using the Auto-Regressive with eXogenous input (ARX) model and Liljencrants-Fant (LF) model with an iteration-based estimation approach. However, the all-pole autoregressive model in the modeling of vocal tract filters cannot provide the locations of anti-formants (zeros), which increases the estimation errors in certain classes of speech sounds, such as nasal, fricative, and stop consonants. In this paper, we propose the Auto-Regressive Moving Average eXogenous with LF (ARMAX-LF) model to extend the ARX-LF model to a wider variety of speech sounds, including vowels and nasalized consonants. The LF model represents the glottal source derivative as a parametrized time-domain model, and the ARMAX model represents the vocal tract as a pole-zero filter with an additional exogenous LF excitation as input. To estimate multiple parameters with fewer errors, we first utilize the powerful nonlinear fitting ability of deep neural networks (DNNs) to build a mapping from extracted glottal source derivatives or speech waveforms to corresponding LF parameters. Then, glottal source and vocal tract parameters can be estimated with fewer estimation errors and without any iterations as in the analysis-by-synthesis strategy. Experimental results with synthesized speech using the linear source-filter model, synthesized speech using the physical model, and real speech signals showed that the proposed ARMAX-LF model with a DNN-based estimation method can estimate the parameters of both vowels and nasalized sounds with fewer errors and estimation time.

语音建模深度学习声学参数语音合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。