arXiv:2511.17477cs.SDcs.AI2025-11

用多模态模型精准识别阿拉伯语发音错误,助力古兰经诵读学习。

Enhancing Quranic Learning: A Multimodal Deep Learning Approach for Arabic Phoneme Recognition

  • 融合语音与文本特征,用Transformer架构统一建模
  • 在29个阿拉伯音素上达到高精度,最优融合策略提升效果
  • 适合对古兰经诵读、语音教学感兴趣的开发者和研究者

近年来,多模态深度学习显著提升了语音分析与发音评估系统的能力。准确检测发音错误仍是阿拉伯语中的关键挑战,尤其在古兰经诵读中,细微的语音差异可能改变语义。本文提出一种基于Transformer的多模态框架,用于阿拉伯音素误读检测,结合声学与文本表示以提高精度与鲁棒性。该框架融合UniSpeech生成的声学嵌入与BERT提取的文本嵌入(来自Whisper转录),构建统一表征,捕捉语音细节与语言上下文。为探索最佳融合方式,测试了早期、中期与晚期融合策略,在包含29个阿拉伯音素(含8种哈菲兹音)的两个数据集上进行评估,涉及11位母语者。额外引入公开YouTube录音增强数据多样性与泛化能力。使用准确率、精确率、召回率与F1分数评估性能,结果表明UniSpeech-BERT多模态配置表现优异,融合型Transformer架构在音素级误读检测中有效。研究推动了智能、说话人无关、多模态计算机辅助语言学习系统的发展,为技术赋能的古兰经发音训练及更广泛的语音教育应用提供可行路径。

原文摘要 · Abstract (English)

Recent advances in multimodal deep learning have greatly enhanced the capability of systems for speech analysis and pronunciation assessment. Accurate pronunciation detection remains a key challenge in Arabic, particularly in the context of Quranic recitation, where subtle phonetic differences can alter meaning. Addressing this challenge, the present study proposes a transformer-based multimodal framework for Arabic phoneme mispronunciation detection that combines acoustic and textual representations to achieve higher precision and robustness. The framework integrates UniSpeech-derived acoustic embeddings with BERT-based textual embeddings extracted from Whisper transcriptions, creating a unified representation that captures both phonetic detail and linguistic context. To determine the most effective integration strategy, early, intermediate, and late fusion methods were implemented and evaluated on two datasets containing 29 Arabic phonemes, including eight hafiz sounds, articulated by 11 native speakers. Additional speech samples collected from publicly available YouTube recordings were incorporated to enhance data diversity and generalization. Model performance was assessed using standard evaluation metrics: accuracy, precision, recall, and F1-score, allowing a detailed comparison of the fusion strategies. Experimental findings show that the UniSpeech-BERT multimodal configuration provides strong results and that fusion-based transformer architectures are effective for phoneme-level mispronunciation detection. The study contributes to the development of intelligent, speaker-independent, and multimodal Computer-Aided Language Learning (CALL) systems, offering a practical step toward technology-supported Quranic pronunciation training and broader speech-based educational applications.

语音识别多模态古兰经学习深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。