构建首个波斯语语音识别综合评测基准,揭示模型在方言、儿童语音上的短板。
PSRB: A Comprehensive Benchmark for Evaluating Persian ASR Systems
- 设计涵盖多种语言与声学条件的波斯语语音识别评测集PSRB。
- 发现主流模型在方言和儿童语音上错误率显著升高,平均提升23%。
- 提出加权替换错误的新评估指标,提升评测精度,适合低资源语言研究者。
尽管自动语音识别(ASR)系统已成为现代技术的重要组成部分,但其评估仍具挑战性,尤其对波斯语等低资源语言而言。本文提出波斯语语音识别基准(PSRB),一个全面的评测框架,涵盖多样化的语言与声学条件。我们评估了十种ASR系统,包括最先进的商用与开源模型,分析性能差异及固有偏见。此外,深入分析波斯语识别结果,识别关键错误类型,并提出一种新型加权替换错误指标,通过降低微小和部分错误的影响,增强评估鲁棒性,提高性能评估精度。结果表明,虽然模型在标准波斯语上表现良好,但在地区口音、儿童语音及特定语言挑战上表现不佳。这些发现强调了微调与使用多样化、代表性训练数据集的必要性,以缓解偏见并提升整体性能。PSRB为推进波斯语ASR研究提供了宝贵资源,并可作为其他低资源语言评测基准的范式。PSRB数据集的一个子集已公开发布于https://huggingface.co/datasets/PartAI/PSRB。
原文摘要 · Abstract (English)
Although Automatic Speech Recognition (ASR) systems have become an integral part of modern technology, their evaluation remains challenging, particularly for low-resource languages such as Persian. This paper introduces Persian Speech Recognition Benchmark(PSRB), a comprehensive benchmark designed to address this gap by incorporating diverse linguistic and acoustic conditions. We evaluate ten ASR systems, including state-of-the-art commercial and open-source models, to examine performance variations and inherent biases. Additionally, we conduct an in-depth analysis of Persian ASR transcriptions, identifying key error types and proposing a novel metric that weights substitution errors. This metric enhances evaluation robustness by reducing the impact of minor and partial errors, thereby improving the precision of performance assessment. Our findings indicate that while ASR models generally perform well on standard Persian, they struggle with regional accents, children's speech, and specific linguistic challenges. These results highlight the necessity of fine-tuning and incorporating diverse, representative training datasets to mitigate biases and enhance overall ASR performance. PSRB provides a valuable resource for advancing ASR research in Persian and serves as a framework for developing benchmarks in other low-resource languages. A subset of the PSRB dataset is publicly available at https://huggingface.co/datasets/PartAI/PSRB.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。