用电磁发音轨迹高效合成高质量语音,参数少4.9倍且速度更快。
Fast, High-Quality and Parameter-Efficient Articulatory Synthesis using Differentiable DSP
- 将电磁发音轨迹与可微信号处理结合,实现低参数语音合成。
- 合成语音词错误率6.67%,主观评分3.74,优于当前最佳模型。
- 仅需0.4M参数即可达到高音质,推理速度比基线快4.9倍。
电磁发音轨迹(EMA)提供了声道滤波器的低维表示,常被用作语音合成中的自然、具物理意义的特征。可微数字信号处理(DDSP)是一种高效的音频合成框架。将低维EMA特征与DDSP结合,可显著提升语音合成的计算效率。本文提出一种快速、高质量且参数高效的DDSP发音体声码器,能从EMA、基频(F0)和响度生成语音。通过引入多种技术解决谐波/噪声失衡问题,并采用多分辨率对抗损失以提升合成质量。模型在转录词错误率(WER)上达到6.67%,平均意见分(MOS)为3.74,较当前最优基线分别提升1.63%和0.16。该声码器在CPU推理时比基线快4.9倍,仅需0.4M参数即可生成与当前最优模型相当的语音质量,而后者需9M参数。
原文摘要 · Abstract (English)
Articulatory trajectories like electromagnetic articulography (EMA) provide a low-dimensional representation of the vocal tract filter and have been used as natural, grounded features for speech synthesis. Differentiable digital signal processing (DDSP) is a parameter-efficient framework for audio synthesis. Therefore, integrating low-dimensional EMA features with DDSP can significantly enhance the computational efficiency of speech synthesis. In this paper, we propose a fast, high-quality, and parameter-efficient DDSP articulatory vocoder that can synthesize speech from EMA, F0, and loudness. We incorporate several techniques to solve the harmonics / noise imbalance problem, and add a multi-resolution adversarial loss for better synthesis quality. Our model achieves a transcription word error rate (WER) of 6.67% and a mean opinion score (MOS) of 3.74, with an improvement of 1.63% and 0.16 compared to the state-of-the-art (SOTA) baseline. Our DDSP vocoder is 4.9x faster than the baseline on CPU during inference, and can generate speech of comparable quality with only 0.4M parameters, in contrast to the 9M parameters required by the SOTA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。