arXiv:2502.11572eess.AScs.SD2025-02中稿 · IEEE SLT 2024被引 3

用670小时数据微调Whisper,让其更准识别罕见词。

Improving Rare-Word Recognition of Whisper in Zero-Shot Settings

  • 用少量共用语音数据微调Whisper,实现上下文引导。
  • 在11个数据集上罕见词识别率提升45.6%。
  • 能力可泛化到未见语言,适合多语言场景。

Whisper虽基于68万小时网络音频训练,仍难以识别领域术语等罕见词,现有方法依赖提示工程进行上下文引导。本文提出一种监督学习策略,仅用670小时Common Voice英语数据微调Whisper,实现上下文引导指令。实验表明,该模型在11个开源英语数据集上,罕见词识别率较基线提升45.6%,未见词识别率提升60.8%。令人惊讶的是,其上下文引导能力还可泛化至微调时未见的语言。

原文摘要 · Abstract (English)

Whisper, despite being trained on 680K hours of web-scaled audio data, faces difficulty in recognising rare words like domain-specific terms, with a solution being contextual biasing through prompting. To improve upon this method, in this paper, we propose a supervised learning strategy to fine-tune Whisper for contextual biasing instruction. We demonstrate that by using only 670 hours of Common Voice English set for fine-tuning, our model generalises to 11 diverse open-source English datasets, achieving a 45.6% improvement in recognition of rare words and 60.8% improvement in recognition of words unseen during fine-tuning over the baseline method. Surprisingly, our model's contextual biasing ability generalises even to languages unseen during fine-tuning.

语音识别罕见词零样本微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。