arXiv:2510.24570cs.CL2025-10中稿 · ICASSP 2026被引 3

用无标签数据让Whisper语音识别更懂航空通信

BEST-RQ-Based Self-Supervised Learning for Whisper Domain Adaptation

  • 用BEST-RQ目标+知识蒸馏适配Whisper编码器
  • 仅用5000小时无字幕语音,性能比微调模型高12%
  • 适合低资源领域语音识别研究者参考

自动语音识别系统在缺乏标注数据的低资源场景下表现不佳。本文提出BEARD(BEST-RQ编码器适应与重训练及蒸馏),一种利用无标签数据适配Whisper编码器的新框架。不同于传统自监督学习方法,BEARD将BEST-RQ目标与冻结教师编码器的知识蒸馏结合,确保编码器与预训练解码器互补。实验聚焦于具有非母语发音、噪声和专业术语的航空管制(ATC)通信领域,使用约5000小时未转录语音进行BEARD训练,仅2小时有标注语音用于微调。该方法显著优于先前基线和微调模型,在相对性能上提升12%。据我们所知,这是首个针对Whisper进行领域自适应的自监督学习工作。

原文摘要 · Abstract (English)

Automatic Speech Recognition (ASR) systems, despite large multilingual training, struggle in low-resource scenarios where labeled data is scarce. We propose BEARD (BEST-RQ Encoder Adaptation with Re-training and Distillation), a novel framework designed to adapt Whisper's encoder with unlabeled data. Unlike traditional self-supervised learning methods, BEARD uniquely combines a BEST-RQ objective with knowledge distillation from a frozen teacher encoder, ensuring the encoder's complementarity with the pre-trained decoder. Our experiments focus on the ATCO2 corpus from the challenging Air Traffic Control (ATC) communications domain, characterized by non-native speech, noise, and specialized phraseology. Using about 5,000 hours of untranscribed speech for BEARD and 2 hours of transcribed speech for fine-tuning, the proposed approach significantly outperforms previous baseline and fine-tuned model, achieving a relative improvement of 12% compared to the fine-tuned model. To the best of our knowledge, this is the first work to use a self-supervised learning objective for domain adaptation of Whisper.

语音识别自监督学习领域适应Whisper

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。