arXiv:2606.24086eess.AScs.SD2026-06中稿 · Interspeech 2026

针对阿拉伯语发音错误检测数据少难题,提出两阶段融合框架提升识别精度。

A Fusion-Aware Two-Stage Framework for Mispronunciation Detection and Diagnosis in Low-Resource Modern Standard Arabic

论文配图:A Fusion-Aware Two-Stage Framework for Mispronunciation Detection and Diagnosis in Low-Resource Modern Standard Arabic
图 1 · 摘自论文原文
  • 用预训练编码器+因果卷积网络保留细微发音差异
  • 两阶段训练:先学通用特征,再适配真实学习者数据
  • 多检查点集成+N元语法重打分,提升预测稳定性

准确的音素识别对现代标准阿拉伯语(MSA)发音错误检测与诊断(MDD)至关重要,但受限于数据稀缺和合成数据与真实数据之间的域差距。本文提出一种两阶段端到端框架,结合预训练编码器与因果膨胀时间卷积网络,以保留精细的语音变化。采用分层两阶段策略:首先在母语者/合成语料上学习通用映射,随后适应稀缺的真实学习者数据,缓解域偏移且避免过度校正。通过多检查点集成推理与N-gram重打分进一步提升预测稳定性。在QuranMB.v2测试集上,系统取得0.7201的F1分数,相比基线(0.4414)提升63.1%,性能位居IqraEval.2挑战赛首位,为低资源MSA的MDD树立新基准。

原文摘要 · Abstract (English)

Accurate phoneme recognition is pivotal for mispronunciation detection and diagnosis (MDD) in modern standard Arabic (MSA), yet remains constrained by data scarcity and the synthetic-real domain gap. This work proposes a two-stage end-to-end framework. It integrates a pre-trained encoder with causal dilated temporal convolutional networks to preserve fine-grained phonetic variations. A hierarchical two-stage strategy first learns general mappings from native/synthetic corpora, then adapts to scarce real learner data to mitigate domain shift without over-correction. Prediction stability is further enhanced via multi-checkpoint ensemble inference with N-gram rescoring. Evaluated on the QuranMB.v2 test set, our system achieves an F1-score of $0.7201$, a $63.1$\% relative improvement over baseline ($0.4414$). This performance ranks at the top of the IqraEval.2 Challenge, establishing a new state-of-the-art for low-resource MSA in MDD.

语音识别发音纠错低资源语言两阶段模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。