arXiv:2509.00094eess.AScs.AI2025-09被引 1

用深度学习自动检测并纠正学经者诵读《古兰经》的发音错误。

Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning

  • 98%自动化流程生成高质量诵经数据集,含音频与标注。
  • 在真实错误数据上达到1.94%音素错误率和75.8%塔吉维德识别率。
  • 自研音标系统编码诵读规则,适合语言教学与语音评估场景。

评估口语发音具有挑战性,而为机器学习模型量化发音指标更难。但对《古兰经》而言,由于穆斯林学者确立的严格诵读规则(塔吉维德),实现高效评估成为可能。尽管如此,高质量标注数据稀缺仍是主要障碍。本文提出:(1) 98%自动化的数据生产流程,涵盖专家诵读收集、基于微调wav2vec2-BERT的停顿点分割、段落转录及通过新提出的Tasmeea算法验证;(2) 848小时音频(286,000个标注语句);(3) qdat_bench基准,覆盖音素、符号标注和塔吉维德规则(格纳、卡尔卡拉、麦德)的真实发音错误,共159个样本;(4) 基于ASR的发音错误检测方法,使用作者自研《古兰经音标脚本》(QPS)编码塔吉维德规则(不同于现代标准阿拉伯语的IPA)。QPS采用11级编码:音素层(阿拉伯字母加长短元音)与属性层(每个音素的发音特征)。我们进一步提出多层级CTC模型,在测试集上实现0.21%平均音素错误率(PER),在qdat_bench上达1.94% PER,塔吉维德F1得分为75.8%。相关成果已开源:https://obadx.github.io/quran-muaalem/en/

原文摘要 · Abstract (English)

Assessing spoken language is challenging, and quantifying pronunciation metrics for machine learning models is even harder. However, for the Holy Quran, this task is enabled by the rigorous recitation rules (Tajweed) established through the efforts of Muslim scholars, making highly effective assessment possible. Despite this advantage, the scarcity of high-quality annotated data remains a significant barrier. In this work, we bridge these gaps by introducing: (1) A 98% automated pipeline to produce high-quality Quranic datasets -- encompassing collection of recitations from expert reciters, segmentation at pause points (waqf) using our fine-tuned wav2vec2-BERT model, transcription of segments, and transcript verification via our novel Tasmeea algorithm; (2) 848 hours of audio (286K annotated utterances); (3) qdat_bench, a benchmark covering phonemes, diacritization, and Tajweed rules (Ghunnah, Qalqalah, Madd) on real recitation errors containing 159 samples; (4) A novel ASR-based approach for pronunciation error detection utilizing our custom Quran Phonetic Script (QPS) to encode Tajweed rules (unlike the IPA standard for Modern Standard Arabic). QPS uses an 11-level script: phoneme level (encoding Arabic letters with short/long vowels) and sifat level (encoding articulation characteristics of every phoneme). We further present comprehensive modeling with our novel multi-level CTC model, which achieved 0.21% and 1.94% average Phoneme Error Rate (PER) on the test set and qdat_bench respectively, with a 75.8% Tajweed F1 score. We release our work as open-source: https://obadx.github.io/quran-muaalem/en/

语音识别古兰经发音纠错深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。