arXiv:2505.17538cs.CLcs.SD2025-05中稿 · Interspeech 2025被引 1

用超大瑞典语语料微调Whisper,显著提升识别准确率。

Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition

  • 基于大规模多变瑞典语数据微调Whisper模型
  • 最佳模型相比原版Whisper-large-v3平均降低47%错误率
  • 适合关注小语种语音识别的开发者与研究者

本文提出一系列针对瑞典语的微调Whisper模型,训练数据规模和多样性在该中等资源语言中前所未有。由于小语种常在多语言训练集中被忽视,通过微调现有多语言模型可显著提升性能。实验显示,所有尺寸模型在瑞典语上的表现均优于OpenAI发布的Whisper。尤其值得注意的是,在FLEURS、Common Voice和NST三个评测集上,最佳模型相较Whisper-large-v3平均降低47%的词错误率(WER)。

原文摘要 · Abstract (English)

This work presents a suite of fine-tuned Whisper models for Swedish, trained on a dataset of unprecedented size and variability for this mid-resourced language. As languages of smaller sizes are often underrepresented in multilingual training datasets, substantial improvements in performance can be achieved by fine-tuning existing multilingual models, as shown in this work. This work reports an overall improvement across model sizes compared to OpenAI's Whisper evaluated on Swedish. Most notably, we report an average 47% reduction in WER comparing our best performing model to OpenAI's whisper-large-v3, in evaluations across FLEURS, Common Voice, and NST.

语音识别Whisper瑞典语微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。