用文本生成语音,让低资源语言语音识别效果提升30%以上。
From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition
- 用现成语音合成模型把文本转成大量合成语音。
- 仅需几十小时真实语音,就生成超50万小时高质量合成语音。
- 适合想低成本提升多语言语音识别的团队使用。
近期自动语音识别(ASR)的进步主要依赖大规模语音语料库。然而,如何覆盖资源有限的多样语言仍面临巨大挑战。本文提出语音反向翻译(Speech Back-Translation),通过现成的文本转语音(TTS)模型将大规模文本语料转化为合成语音,以提升多语言ASR模型性能。我们证明,仅需几十小时的真实标注语音即可训练出高质量的TTS模型,生成的合成语音量可达原始数据的数百倍。为评估合成语音质量,我们构建了基于可理解性的评估框架,并确立了合成数据对ASR训练有益的明确阈值。利用该方法,我们在十种语言中生成超过50万小时的合成语音,继续预训练Whisper-large-v3,实现平均转录错误率降低30%以上。结果表明,该方法在提升多语言ASR系统方面具有显著的可扩展性和有效性。
原文摘要 · Abstract (English)
Recent advances in Automatic Speech Recognition (ASR) have been largely fueled by massive speech corpora. However, extending coverage to diverse languages with limited resources remains a formidable challenge. This paper introduces Speech Back-Translation, a scalable pipeline that improves multilingual ASR models by converting large-scale text corpora into synthetic speech via off-the-shelf text-to-speech (TTS) models. We demonstrate that just tens of hours of real transcribed speech can effectively train TTS models to generate synthetic speech at hundreds of times the original volume while maintaining high quality. To evaluate synthetic speech quality, we develop an intelligibility-based assessment framework and establish clear thresholds for when synthetic data benefits ASR training. Using Speech Back-Translation, we generate more than 500,000 hours of synthetic speech in ten languages and continue pre-training Whisper-large-v3, achieving average transcription error reductions of over 30\%. These results highlight the scalability and effectiveness of Speech Back-Translation for enhancing multilingual ASR systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。