用合成语音自迭代优化语音识别模型,显著降低错误率。
A Self-Refining Framework for Enhancing ASR Using TTS-Synthesized Data
- 用现有模型生成伪标签,训练高保真语音合成系统,形成闭环优化。
- 在6000小时无标注语音上,普通话错误率降20%,中英混杂降50%。
- 适合资源少或领域特定的语音识别场景,无需大量标注数据。
我们提出一种自精炼框架,仅使用无标注数据提升语音识别(ASR)性能。流程从现有ASR模型对未标注语音生成伪标签开始,据此训练高保真文本到语音(TTS)系统;随后将合成的语音-文本对回注入原始ASR系统,完成闭环自我改进。我们在台湾闽南语语音上验证了该框架的有效性。利用6000小时无标注语音、适量文本数据及AI生成合成内容,我们将Whisper-large-v2适配为专用模型Twister。相比Whisper,Twister在普通话任务上错误率降低最高达20%,在中英混杂语音任务上降低达50%。结果表明,该框架是伪标签自蒸馏方法的有力替代方案,为低资源或领域特定场景下的ASR性能提升提供了实用路径。
原文摘要 · Abstract (English)
We propose a self-refining framework that enhances ASR performance with only unlabeled datasets. The process starts with an existing ASR model generating pseudo-labels on unannotated speech, which are then used to train a high-fidelity text-to-speech (TTS) system. Then, synthesized speech text pairs are bootstrapped into the original ASR system, completing the closed-loop self-improvement cycle. We demonstrated the effectiveness of the framework on Taiwanese Mandarin speech. Leveraging 6,000 hours of unlabeled speech, a moderate amount of text data, and synthetic content from the AI models, we adapt Whisper-large-v2 into a specialized model, Twister. Twister reduces error rates by up to 20% on Mandarin and 50% on Mandarin-English code-switching benchmarks compared to Whisper. Results highlight the framework as a compelling alternative to pseudo-labeling self-distillation approaches and provides a practical pathway for improving ASR performance in low-resource or domain-specific settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。