用廉价弱标签训练语音模型,仅需少量精准数据即可达到高精度。
From Weak Labels to Strong Results: Utilizing 5,000 Hours of Noisy Classroom Transcripts with Minimal Accurate Data
- 先用5000小时噪声转录文本监督预训练,再微调真实标注数据。
- 在真实课堂场景中,相比其他方法错误率降低12.3%。
- 适合数据标注成本高的教育类语音识别任务。
近年来语音识别的进步依赖于海量标注数据的训练。然而,课堂自动语音识别(ASR)面临真实世界挑战:拥有大量廉价的弱转录文本,但仅有少量准确的黄金标准数据。在此低资源环境下,重新标注的成本过高,难以实施。为此,我们提出弱监督预训练(WSP),一种两阶段方法:模型首先在弱转录文本上以监督方式预训练,随后在高质量数据上微调。基于合成与真实弱转录文本的实验表明,WSP优于其他方法,确立其在真实场景下低资源语音识别中的有效性。
原文摘要 · Abstract (English)
Recent progress in speech recognition has relied on models trained on vast amounts of labeled data. However, classroom Automatic Speech Recognition (ASR) faces the real-world challenge of abundant weak transcripts paired with only a small amount of accurate, gold-standard data. In such low-resource settings, high transcription costs make re-transcription impractical. To address this, we ask: what is the best approach when abundant inexpensive weak transcripts coexist with limited gold-standard data, as is the case for classroom speech data? We propose Weakly Supervised Pretraining (WSP), a two-step process where models are first pretrained on weak transcripts in a supervised manner, and then fine-tuned on accurate data. Our results, based on both synthetic and real weak transcripts, show that WSP outperforms alternative methods, establishing it as an effective training methodology for low-resource ASR in real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。