arXiv:2505.24656cs.CLcs.SD2025-05被引 2

用伪标签与自监督结合,提升语音识别在低资源语言下的适应能力

MSDA: Combining Pseudo-labeling and Self-Supervision for Unsupervised Domain Adaptation in ASR

  • 分两阶段融合自监督与半监督学习,提升模型泛化性
  • 在希腊语等低资源语言上显著优于现有方法
  • 适合标注数据少或噪声多的语音识别场景

本文研究了用于自动语音识别(ASR)的元伪标签(Meta PL)无监督域适应框架。提出一种多阶段域适应流程(MSDA),是一种样本高效的两阶段适配方法,结合自监督学习与半监督技术。MSDA旨在增强ASR模型的鲁棒性和泛化能力,使其更适应多样环境,尤其适用于希腊语等低资源语言及标注数据稀缺或噪声较多的弱监督场景。大量实验表明,Meta PL可有效应用于ASR任务,达到当前最优性能,显著超越现有先进方法,并为无监督域适应提供更稳健的解决方案。消融实验凸显了在结合自监督与自训练时采用级联策略的必要性。

原文摘要 · Abstract (English)

In this work, we investigate the Meta PL unsupervised domain adaptation framework for Automatic Speech Recognition (ASR). We introduce a Multi-Stage Domain Adaptation pipeline (MSDA), a sample-efficient, two-stage adaptation approach that integrates self-supervised learning with semi-supervised techniques. MSDA is designed to enhance the robustness and generalization of ASR models, making them more adaptable to diverse conditions. It is particularly effective for low-resource languages like Greek and in weakly supervised scenarios where labeled data is scarce or noisy. Through extensive experiments, we demonstrate that Meta PL can be applied effectively to ASR tasks, achieving state-of-the-art results, significantly outperforming state-of-the-art methods, and providing more robust solutions for unsupervised domain adaptation in ASR. Our ablations highlight the necessity of utilizing a cascading approach when combining self-supervision with self-training.

语音识别域适应自监督低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。