arXiv:2506.07722cs.SDeess.AS2025-06中稿 · terspeech 2025 and…被引 13

构建首个阿拉伯语发音评估基准,以古兰经诵读为案例

Towards a Unified Benchmark for Arabic Pronunciation Assessment: Quranic Recitation as Case Study

  • 基于古兰经诵读构建统一评估框架
  • 提出首个公开的阿拉伯语发音错误检测测试集
  • 为阿拉伯语语音技术研究提供标准化工具

我们提出一个统一的现代标准阿拉伯语(MSA)发音错误检测基准,以古兰经诵读为案例。该方法建立了涵盖数据处理、专用于MSA发音特点的音素集设计,以及首个公开可用的测试集——古兰经发音错误基准(QuranMB.v1)的完整流程。同时,我们评估了多个基线模型,提供了初步性能分析,揭示了评估MSA发音的潜力与挑战。通过建立这一标准化框架,旨在推动阿拉伯语语音技术及相关应用的研究与发展。

原文摘要 · Abstract (English)

We present a unified benchmark for mispronunciation detection in Modern Standard Arabic (MSA) using Qur'anic recitation as a case study. Our approach lays the groundwork for advancing Arabic pronunciation assessment by providing a comprehensive pipeline that spans data processing, the development of a specialized phoneme set tailored to the nuances of MSA pronunciation, and the creation of the first publicly available test set for this task, which we term as the Qur'anic Mispronunciation Benchmark (QuranMB.v1). Furthermore, we evaluate several baseline models to provide initial performance insights, thereby highlighting both the promise and the challenges inherent in assessing MSA pronunciation. By establishing this standardized framework, we aim to foster further research and development in pronunciation assessment in Arabic language technology and related applications.

语音评估阿拉伯语古兰经音素建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。