arXiv:2503.08954eess.AScs.CL2025-03被引 6

用流模型生成语音增强数据,显著降低语音识别错误率。

An Exhaustive Evaluation of TTS- and VC-based Data Augmentation for ASR

  • 采用基于流的TTS/VC模型生成多样化合成语音。
  • 联合增强多种语音属性,使错误率相对降低11%至35%。
  • 适合需要提升低资源语音识别性能的研究者。

近年来,使用文本转语音(TTS)或语音转换(VC)生成合成数据来扩充自动语音识别(ASR)系统的训练数据越来越流行。尽管已有研究证明该方法可提升ASR性能,但由于合成语音多样性不足,简单混合真实与合成数据往往无法取得最佳效果。本文利用最近提出的基于流的TTS/VC模型以提升语音多样性,并评估了不同语音属性增强对多个ASR模型词错误率(WER)的影响。结果表明,音高增强和基于VC的说话人增强在本设置中无效。而联合增强其余所有语音属性,使Conformer-Transducer模型在Common Voice上相对错误率降低11%,在LibriSpeech上最高降低35%。

原文摘要 · Abstract (English)

Augmenting the training data of automatic speech recognition (ASR) systems with synthetic data generated by text-to-speech (TTS) or voice conversion (VC) has gained popularity in recent years. Several works have demonstrated improvements in ASR performance using this augmentation approach. However, because of the lower diversity of synthetic speech, naively combining synthetic and real data often does not yield the best results. In this work, we leverage recently proposed flow-based TTS/VC models allowing greater speech diversity, and assess the respective impact of augmenting various speech attributes on the word error rate (WER) achieved by several ASR models. Pitch augmentation and VC-based speaker augmentation are found to be ineffective in our setup. Jointly augmenting all other attributes reduces the WER of a Conformer-Transducer model by 11\% relative on Common Voice and by up to 35\% relative on LibriSpeech compared to training on real data only.

语音识别数据增强流模型合成语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。