arXiv:2607.05224cs.CL2026-07

用迭代伪标签提升中英混用语音识别,显著降低错误率。

Progressive Refinement: An Iterative Pseudo-Labeling Approach for Mandarin-English Code-Switching ASR

论文配图:Progressive Refinement: An Iterative Pseudo-Labeling Approach for Mandarin-English Code-Switching ASR
图 1 · 摘自论文原文
  • 通过无标签数据生成伪标签,构建半监督训练集。
  • 在SEAME数据集上,混合错误率分别降低6.35%和8.29%。
  • 适合缺乏标注数据的多语言语音识别研究者。

中英混用(CS)语音识别因标注数据稀缺面临挑战。本文首次提出一种迭代伪标签训练方法,有效利用无标签数据提升模型性能。流程包括:从大规模无标签语料生成伪标签,构建半监督数据集;采用两阶段双语模型训练(预训练+微调);通过迭代优化增强复杂混用场景的识别能力。实验表明,该方法在SEAME数据集的devman子集(6.35%)和devsge子集(8.29%)上实现显著的混合错误率下降,推动了中文-英文混用语音识别系统的发展。

原文摘要 · Abstract (English)

Code-switching (CS), alternating languages within the same utterance, poses significant challenges for automatic speech recognition (ASR) due to limited CS training data. This paper applies an iterative pseudo-labeling training approach to CS-ASR for the first time, demonstrating its effectiveness in leveraging unlabeled data to improve CS-ASR performance. The approach comprises three phases: pseudo-label generation, two-stage bilingual model training, and iterative improvements. It begins by generating pseudo-labels from a large unlabeled corpus, creating a semi-supervised dataset. This dataset supports a two-stage training framework where the model is pre-trained and then fine-tuned on supervised CS data. Iterative refinements further enhance the model's accuracy in handling complex CS scenarios. Our approach significantly advances CS-ASR systems, achieving notable Mix Error Rate (MER) reductions on SEAME's devman (6.35%) and devsge (8.29%) subsets.

语音识别混用识别伪标签半监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。