arXiv:2410.12279eess.AScs.AI2024-10

对比扩散模型与均方误差,发现前者在语音合成中更适配ASR训练。

Beyond Oversmoothing: Evaluating DDPM and MSE for Scalable Speech Synthesis in ASR

论文配图:Beyond Oversmoothing: Evaluating DDPM and MSE for Scalable Speech Synthesis in ASR
图 1 · 摘自论文原文
  • 用扩散模型替代传统均方误差训练语音合成器。
  • 数据量和语者多样性提升时,扩散模型性能更优,达到1.46的最优真实/合成语音误识率比。
  • 适合关注语音合成如何服务自动语音识别的科研人员。

合成语音已接近人类自然度,但以人耳评判自然的语音合成数据训练的自动语音识别(ASR)系统,在真实语音上表现依然不佳。本文探究这一现象是否源于语音合成模型常见的过平滑问题,重点考察在扩大语音合成训练数据规模时,面向ASR的语音合成模型表现。系统比较了基于去噪扩散概率模型(DDPM)与均方误差(MSE)的语音合成方法在ASR训练中的表现。通过调整训练时长(小时数)和不同说话人数量,评估两种方法的可扩展性。结果表明,在相同模型规模下,DDPM能更有效地利用更多数据和更丰富的说话人多样性;最终实现了目前最佳的真实与合成语音误识率比(1.46),但仍存在显著差距。

原文摘要 · Abstract (English)

Synthetically generated speech has rapidly approached human levels of naturalness. However, the paradox remains that ASR systems, when trained on TTS output that is judged as natural by humans, continue to perform badly on real speech. In this work, we explore whether this phenomenon is due to the oversmoothing behaviour of models commonly used in TTS, with a particular focus on the behaviour of TTS-for-ASR as the amount of TTS training data is scaled up. We systematically compare Denoising Diffusion Probabilistic Models (DDPM) to Mean Squared Error (MSE) based models for TTS, when used for ASR model training. We test the scalability of the two approaches, varying both the number hours, and the number of different speakers. We find that for a given model size, DDPM can make better use of more data, and a more diverse set of speakers, than MSE models. We achieve the best reported ratio between real and synthetic speech WER to date (1.46), but also find that a large gap remains.

语音合成扩散模型ASR训练可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。