arXiv:2506.08836cs.CLcs.HC2025-06被引 3

为瑞士德语构建300小时真实场景语音数据集,提升低资源语言语音识别效果

Advancing STT for Low-Resource Real-World Speech

  • 构建300小时瑞士方言真实录音数据集SRB-300,覆盖多种口语环境
  • 微调Whisper模型后词错误率降低19%~33%,BLEU提升8%~40%
  • 适合研究低资源语言、方言语音识别及真实场景下STT系统优化

瑞士德语是一种资源稀缺的语言,存在多种差异显著的方言,无统一书写形式,通常需转写为标准德语。现有数据集多来自受控环境,虽训练出高效语音转文本(STT)模型,但对自然对话语音表现不佳。本文提出新的SRB-300数据集,包含39个瑞士广播电视台采集的300小时长音频,涵盖主要方言与真实环境下的自发性口语。该数据集突破了以往句子级数据的局限。我们在SRB-300上微调多个OpenAI Whisper模型,性能显著优于零样本基线:词错误率(WER)降低19%至33%,BLEU分数提升8%至40%。最佳模型large-v3实现WER 17.1%、BLEU 74.8。该成果对开发瑞士德语及其他低资源语言在真实场景中的鲁棒语音识别系统具有重要意义。

原文摘要 · Abstract (English)

Swiss German is a low-resource language represented by diverse dialects that differ significantly from Standard German and from each other, lacking a standardized written form. As a result, transcribing Swiss German involves translating into Standard German. Existing datasets have been collected in controlled environments, yielding effective speech-to-text (STT) models, but these models struggle with spontaneous conversational speech. This paper, therefore, introduces the new SRB-300 dataset, a 300-hour annotated speech corpus featuring real-world long-audio recordings from 39 Swiss German radio and TV stations. It captures spontaneous speech across all major Swiss dialects recorded in various realistic environments and overcomes the limitation of prior sentence-level corpora. We fine-tuned multiple OpenAI Whisper models on the SRB-300 dataset, achieving notable enhancements over previous zero-shot performance metrics. Improvements in word error rate (WER) ranged from 19% to 33%, while BLEU scores increased between 8% and 40%. The best fine-tuned model, large-v3, achieved a WER of 17.1% and a BLEU score of 74.8. This advancement is crucial for developing effective and robust STT systems for Swiss German and other low-resource languages in real-world contexts.

语音识别低资源语言方言识别Whisper

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。