arXiv:2506.07659eess.AS2025-06中稿 · Interspeech 2025被引 2

开源全流程语音识别半监督训练框架,支持多语言扩展

Unified Semi-Supervised Pipeline for Automatic Speech Recognition

  • 构建从无标签数据收集到模型训练的完整开源流水线
  • 新伪标签算法TopIPL在多语言上降低18%-40%错误率
  • 适合低资源语言研究者快速搭建半监督语音系统

自动语音识别长期面临标注数据稀缺问题,尽管已有大量研究致力于提升半监督学习算法,但多数工作局限于现有数据集,缺乏面向新数据集或新语言的大规模半监督训练公开框架。本文提出一个全开源的半监督训练框架,覆盖从无标签数据采集、伪标签生成到模型训练的完整流程。该方法可利用公共领域内符合创意共享协议的语音数据,实现任意语言的可扩展数据构建。我们提出一种新型伪标签算法TopIPL,并在低资源语言(葡萄牙语、亚美尼亚语)和高资源语言(西班牙语)中进行评估。结果显示,TopIPL在葡萄牙语上实现18-40%相对词错误率改善,亚美尼亚语为5-16%,西班牙语为2-8%。

原文摘要 · Abstract (English)

Automatic Speech Recognition has been a longstanding research area, with substantial efforts dedicated to integrating semi-supervised learning due to the scarcity of labeled datasets. However, most prior work has focused on improving learning algorithms using existing datasets, without providing a complete public framework for large-scale semi-supervised training across new datasets or languages. In this work, we introduce a fully open-source semi-supervised training framework encompassing the entire pipeline: from unlabeled data collection to pseudo-labeling and model training. Our approach enables scalable dataset creation for any language using publicly available speech data under Creative Commons licenses. We also propose a novel pseudo-labeling algorithm, TopIPL, and evaluate it in both low-resource (Portuguese, Armenian) and high-resource (Spanish) settings. Notably, TopIPL achieves relative WER improvements of 18-40% for Portuguese, 5-16% for Armenian, and 2-8% for Spanish.

语音识别半监督学习开源框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。