arXiv:2506.04981cs.CLcs.SD2025-06中稿 · Interspeech 2025, …被引 4

用相关领域数据和过滤伪标签,提升小样本场景下的语音识别性能。

Better Semi-supervised Learning for Multi-domain ASR Through Incremental Retraining and Data Filtering

  • 分步微调+多模型共识或命名实体识别筛选伪标签。
  • 在Wow和Fisher数据集上相对单步微调提升超20%。
  • 适合资源稀缺领域语音识别,尤其注重效果与效率平衡。

当目标领域标注数据稀少时,微调预训练语音识别模型极具挑战。但通常可获取相关领域的未标注音频和少量标注数据。本文提出一种增量式半监督学习流程:首先融合少量本域标注数据与一个密切相关的辅助数据集,相较无辅助数据提升4%相对性能;随后基于多模型一致性或命名实体识别(NER)对伪标签进行筛选并迭代优化,相比随机选择表现出更慢的性能饱和趋势。在多领域哇呼叫中心(Wow)和Fisher英语语料库上的实验表明,该方法优于单步微调。基于一致性的过滤表现最佳,在Wow上实现22.3%、Fisher上24.8%的相对提升;而NER方法次之,兼具良好性能与更低计算开销。

原文摘要 · Abstract (English)

Fine-tuning pretrained ASR models for specific domains is challenging when labeled data is scarce. But unlabeled audio and labeled data from related domains are often available. We propose an incremental semi-supervised learning pipeline that first integrates a small in-domain labeled set and an auxiliary dataset from a closely related domain, achieving a relative improvement of 4% over no auxiliary data. Filtering based on multi-model consensus or named entity recognition (NER) is then applied to select and iteratively refine pseudo-labels, showing slower performance saturation compared to random selection. Evaluated on the multi-domain Wow call center and Fisher English corpora, it outperforms single-step fine-tuning. Consensus-based filtering outperforms other methods, providing up to 22.3% relative improvement on Wow and 24.8% on Fisher over single-step fine-tuning with random selection. NER is the second-best filter, providing competitive performance at a lower computational cost.

语音识别半监督增量学习数据过滤

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。