arXiv:2606.11429eess.AScs.CL2026-06中稿 · Interspeech 2026

自动选择模型层,让Whisper在低资源语音上表现更好

Gumbel-BEARD: Automatic Layer Selection for Self-Supervised Adaptation of Whisper in Low-Resource Domains

论文配图:Gumbel-BEARD: Automatic Layer Selection for Self-Supervised Adaptation of Whisper in Low-Resource Domains
图 1 · 摘自论文原文
  • 用可训练的硬Gumbel-Softmax自动选Whisper编码器层
  • 10小时标注数据就达到全监督133小时的效果
  • 适合低资源语音场景,尤其儿童和方言数据

语音基础模型在低资源领域常因领域不匹配和数据稀缺而表现不佳。我们提出Gumbel-BEARD,一种端到端可训练的硬Gumbel-Softmax选择器,自动优化Whisper编码器层选择,实现无需人工调参的自监督适配。通过BEST-RQ目标动态适应目标语音特征。在MyST儿童语音语料库上的实验表明:仅用10小时标注数据微调,性能即媲美使用完整133小时标注数据的全监督基线。在MyST上以Whisper-medium取得8.21%的词错误率(WER),在OGI自发语音数据集上以Whisper-small取得11.06%的WER,均创当前最优。在CORAAL上的评估进一步验证其对成人方言域偏移的鲁棒性,相对词错误率降低最高达6%,证明方法在多样低资源条件下的泛化能力。

原文摘要 · Abstract (English)

Speech foundation models often struggle in low-resource domains due to domain mismatch and data scarcity. We propose Gumbel-BEARD, a domain adaptation framework that automates Whisper encoder layer selection via an end-to-end trainable hard Gumbel-Softmax selector. It enables self-supervised adaptation with a BEST-RQ objective that dynamically adapts to target acoustic characteristics without manual tuning. Experiments on the MyST child speech corpus demonstrate efficiency and scalability: with 10 h of labeled data for fine-tuning, our method matches a fully supervised baseline trained on the complete 133 h labeled set. We establish new state-of-the-art word error rates (WERs) of 8.21% using Whisper-medium on MyST and 11.06% using Whisper-small on the OGI Spontaneous dataset. Evaluation on CORAAL further confirms robustness to adult dialectal domain shifts, with up to 6% relative WER reduction, highlighting the generalizability of our approach to diverse low-resource conditions.

语音识别自监督学习Whisper低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。