首个乌尔都语口语语音识别基准测试,揭示模型在真实对话中的表现差异。
WER We Stand: Benchmarking Urdu ASR Models
- 对比Whisper、MMS、Seamless-M4T三类模型在读诵与对话语音上的表现
- Seamless-large在读诵数据上WER最低,Whisper-large在对话数据上最优
- 首次构建乌尔都语对话语音数据集,强调文本规范化对低资源语言的重要性
本文对乌尔都语自动语音识别(ASR)模型进行了全面评估。我们分析了三种模型家族:Whisper、MMS 和 Seamless-M4T,在读诵语音和对话语音两类数据集上的词错误率(WER)。值得注意的是,我们首次提出了专用于乌尔都语ASR模型基准测试的对话语音数据集。结果显示,Seamless-large在读诵语音数据集上表现最佳,而Whisper-large在对话语音数据集上表现最优。此外,本研究指出仅依赖量化指标评估低资源语言如乌尔都语的ASR模型存在局限性,强调建立稳健的乌尔都语文本规范化系统的重要性。研究结果为开发适用于低资源语言的鲁棒语音识别系统提供了宝贵洞见。
原文摘要 · Abstract (English)
This paper presents a comprehensive evaluation of Urdu Automatic Speech Recognition (ASR) models. We analyze the performance of three ASR model families: Whisper, MMS, and Seamless-M4T using Word Error Rate (WER), along with a detailed examination of the most frequent wrong words and error types including insertions, deletions, and substitutions. Our analysis is conducted using two types of datasets, read speech and conversational speech. Notably, we present the first conversational speech dataset designed for benchmarking Urdu ASR models. We find that seamless-large outperforms other ASR models on the read speech dataset, while whisper-large performs best on the conversational speech dataset. Furthermore, this evaluation highlights the complexities of assessing ASR models for low-resource languages like Urdu using quantitative metrics alone and emphasizes the need for a robust Urdu text normalization system. Our findings contribute valuable insights for developing robust ASR systems for low-resource languages like Urdu.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。