arXiv:2502.17284cs.CLcs.SD2025-02被引 4

微调Whisper模型提升荷兰语儿童、老人及非母语者识别准确率

Improving the Inclusivity of Dutch Speech Recognition by Fine-tuning Whisper on the JASMIN-CGN Corpus

  • 在JASMIN-CGN语料库上微调Whisper,针对不同年龄和语言背景群体优化识别
  • 对非母语儿童的识别错误率降低72%,非母语成人降低67%,老年母语者降低65%
  • 强调为少数群体训练模型的重要性,适合语音识别公平性研究者参考

我们测试并研究了在JASMIN-CGN语料库中,针对儿童、老年人及非母语荷兰语使用者的微调版Whisper模型的语音识别表现差异。主要目标是评估说话人年龄和语言背景对Whisper性能的影响。微调后,模型在不同子群体上的词错误率(WER)表现各异。相比零样本性能,微调显著提升:母语儿童相对降低81%、非母语儿童降低72%、非母语成人降低67%、母语老人降低65%。结果表明,训练语音识别模型如Whisper时,必须纳入儿童、老年人及非母语者等代表性不足的子群体。

原文摘要 · Abstract (English)

We test and study the variation in speech recognition of fine-tuned versions of the Whisper model on child, elderly and non-native Dutch speech from the JASMIN-CGN corpus. Our primary goal is to evaluate how speakers' age and linguistic background influence Whisper's performance. Whisper achieves varying Word Error Rates (WER) when fine-tuned on subpopulations of specific ages and linguistic backgrounds. Fine-tuned performance is remarkably better than zero-shot performance, achieving a relative reduction in WER of 81% for native children, 72% for non-native children, 67% for non-native adults, and 65% for native elderly people. Our findings underscore the importance of training speech recognition models like Whisper on underrepresented subpopulations such as children, the elderly, and non-native speakers.

语音识别Whisper包容性多语种

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。