针对阿拉伯语发音错误检测数据少难题,提出两阶段融合框架提升识别精度。
A Fusion-Aware Two-Stage Framework for Mispronunciation Detection and Diagnosis in Low-Resource Modern Standard Arabic

- 用预训练编码器+因果卷积网络保留细微发音差异
- 两阶段训练:先学通用特征,再适配真实学习者数据
- 多检查点集成+N元语法重打分,提升预测稳定性
准确的音素识别对现代标准阿拉伯语(MSA)发音错误检测与诊断(MDD)至关重要,但受限于数据稀缺和合成数据与真实数据之间的域差距。本文提出一种两阶段端到端框架,结合预训练编码器与因果膨胀时间卷积网络,以保留精细的语音变化。采用分层两阶段策略:首先在母语者/合成语料上学习通用映射,随后适应稀缺的真实学习者数据,缓解域偏移且避免过度校正。通过多检查点集成推理与N-gram重打分进一步提升预测稳定性。在QuranMB.v2测试集上,系统取得0.7201的F1分数,相比基线(0.4414)提升63.1%,性能位居IqraEval.2挑战赛首位,为低资源MSA的MDD树立新基准。
原文摘要 · Abstract (English)
Accurate phoneme recognition is pivotal for mispronunciation detection and diagnosis (MDD) in modern standard Arabic (MSA), yet remains constrained by data scarcity and the synthetic-real domain gap. This work proposes a two-stage end-to-end framework. It integrates a pre-trained encoder with causal dilated temporal convolutional networks to preserve fine-grained phonetic variations. A hierarchical two-stage strategy first learns general mappings from native/synthetic corpora, then adapts to scarce real learner data to mitigate domain shift without over-correction. Prediction stability is further enhanced via multi-checkpoint ensemble inference with N-gram rescoring. Evaluated on the QuranMB.v2 test set, our system achieves an F1-score of $0.7201$, a $63.1$\% relative improvement over baseline ($0.4414$). This performance ranks at the top of the IqraEval.2 Challenge, establishing a new state-of-the-art for low-resource MSA in MDD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。