arXiv:2508.21631eess.AS2025-08被引 3

用文本生成语音数据,让语音识别模型性能逼近真实数据训练效果。

Towards Improved Speech Recognition through Optimized Synthetic Data Generation

  • 用先进语音合成模型从纯文本生成语音,实现语音识别数据替代。
  • 优化生成流程后,合成数据训练的识别系统在法语口语数据上表现显著提升。
  • 适合缺乏标注语音数据但有文本资源的研究者或企业使用。

语音识别模型的监督训练依赖带字幕的音频数据,但常因保密问题无法获取。本文提出从纯文本语料库出发,利用具备声音克隆能力的先进文本转语音模型生成合成语音,目标是使基于合成数据训练的自动语音识别(ASR)系统性能接近真实数据训练水平。通过微调、过滤与评估优化合成数据生成过程,并用于训练端到端编码器-解码器式ASR模型。实验在两个魁北克法语的自然对话语音数据集上进行,结果表明,优化生成流程能显著提升合成数据训练所得的ASR系统性能。

原文摘要 · Abstract (English)

Supervised training of speech recognition models requires access to transcribed audio data, which often is not possible due to confidentiality issues. Our approach to this problem is to generate synthetic audio from a text-only corpus using a state-of-the-art text-to-speech model with voice cloning capabilities. Our goal is to achieve automatic speech recognition (ASR) performance comparable to models trained on real data. We explore ways to optimize synthetic data generation through finetuning, filtering and evaluation, and its use for training an end-to-end encoder-decoder ASR model. Experiments were conducted using two datasets of spontaneous, conversational speech in Québec French. We show that improving data generation leads to large improvements in the final ASR system trained on synthetic data.

语音识别合成数据TTS端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。