用深度学习自动评估古兰经诵读的三种发音规则,提升自学效率。
Evaluation of the Pronunciation of Tajweed Rules Based on DNN as a Step Towards Interactive Recitation Learning
- 基于高效网络和注意力机制,从音频中识别三种诵读规则
- 在1500+录音上达到95%以上准确率,最高达99.34%
- 适合想自主练习古兰经诵读的学生与教育系统开发者
正确遵循古兰经诵读规则(Tajweed)对避免诵读错误至关重要,但传统教学受限于合格导师稀缺和时间不足。自动评估系统可提供即时反馈,支持自主练习。本研究基于公开的QDAT数据集(含1500余段音频),使用归一化梅尔频谱图作为输入,采用增强的EfficientNet-B0模型(加入Squeeze-and-Excitation注意力机制)对三种Tajweed规则——延长音(Al Mad)、紧闭鼻音(Ghunnah)、隐藏音(Ikhfaa)进行分类。模型在三类规则上的准确率分别为95.35%、99.34%和97.01%。学习曲线分析表明模型稳健且无过拟合。该方法效率高,为构建互动式Tajweed学习系统奠定基础。
原文摘要 · Abstract (English)
Proper recitation of the Quran, adhering to the rules of Tajweed, is crucial for preventing mistakes during recitation and requires significant effort to master. Traditional methods of teaching these rules are limited by the availability of qualified instructors and time constraints. Automatic evaluation of recitation can address these challenges by providing prompt feedback and supporting independent practice. This study focuses on developing a deep learning model to classify three Tajweed rules - separate stretching (Al Mad), tight noon (Ghunnah), and hide (Ikhfaa) - using the publicly available QDAT dataset, which contains over 1,500 audio recordings. The input data consisted of audio recordings from this dataset, transformed into normalized mel-spectrograms. For classification, the EfficientNet-B0 architecture was used, enhanced with a Squeeze-and-Excitation attention mechanism. The developed model achieved accuracy rates of 95.35%, 99.34%, and 97.01% for the respective rules. An analysis of the learning curves confirmed the model's robustness and absence of overfitting. The proposed approach demonstrates high efficiency and paves the way for developing interactive educational systems for Tajweed study.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。