arXiv:2508.09957cs.CL2025-08被引 3

对比Wav2Vec2与Whisper在库尔德语巴迪尼方言语音识别中的表现

Which one Performs Better? Wav2Vec or Whisper? Applying both in Badini Kurdish Speech to Text (BKSTT)

  • 用儿童故事录音构建巴迪尼方言语音数据集
  • Wav2Vec2模型读写准确率90.38%,显著优于Whisper的65.45%
  • 适合低资源语言语音识别研究者参考

语音转文本(STT)系统应用广泛,但多数语言质量不一。库尔德语中仅有苏莱曼尼方言有可用系统,而巴迪尼、哈瓦拉米等方言尚无。本文针对约两百万使用者的巴迪尼方言,基于8本儿童故事书(共78篇)建立语音数据集,由6名讲述者录制,总时长约17小时。经清洗、分割与分词后,获得近15小时语音数据,含19193个片段和25221个词。采用Wav2Vec2-Large-XLSR-53与Whisper-small模型进行建模。实验显示,基于Wav2Vec2-Large-XLSR-53的转录结果读写准确率达90.38%,可读性为82.67%;而Whisper-small分别为65.45%和53.17%。结果表明,前者在该方言上表现更优。

原文摘要 · Abstract (English)

Speech-to-text (STT) systems have a wide range of applications. They are available in many languages, albeit at different quality levels. Although Kurdish is considered a less-resourced language from a processing perspective, SST is available for some of the Kurdish dialects, for instance, Sorani (Central Kurdish). However, that is not applied to other Kurdish dialects, Badini and Hawrami, for example. This research is an attempt to address this gap. Bandin, approximately, has two million speakers, and STT systems can help their community use mobile and computer-based technologies while giving their dialect more global visibility. We aim to create a language model based on Badini's speech and evaluate its performance. To cover a conversational aspect, have a proper confidence level of grammatical accuracy, and ready transcriptions, we chose Badini kids' stories, eight books including 78 stories, as the textual input. Six narrators narrated the books, which resulted in approximately 17 hours of recording. We cleaned, segmented, and tokenized the input. The preprocessing produced nearly 15 hours of speech, including 19193 segments and 25221 words. We used Wav2Vec2-Large-XLSR-53 and Whisper-small to develop the language models. The experiments indicate that the transcriptions process based on the Wav2Vec2-Large-XLSR-53 model provides a significantly more accurate and readable output than the Whisper-small model, with 90.38% and 65.45% readability, and 82.67% and 53.17% accuracy, respectively.

语音识别低资源语言Wav2Vec2Whisper

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。