arXiv:2606.05569cs.CLcs.SD2026-06中稿 · Interspeech 2026

用语言特定统计图检测跨母语者的发音错误,提升诊断精度。

Domain-Aware Mispronunciation Detection and Diagnosis Using Language-Specific Statistical Graphs

论文配图:Domain-Aware Mispronunciation Detection and Diagnosis Using Language-Specific Statistical Graphs
图 1 · 摘自论文原文
  • 构建音素混淆的有向统计图,建模发音错误模式。
  • 在L2-ARCTIC数据集上达到59.52%的F1分数,优于基线。
  • 适合多语言语音学习系统和发音纠错研究者使用。

近年来,发音错误检测与诊断(MDD)在计算机辅助语言学习和语音技术中变得日益重要。本文提出一种构建统计图的方法,使模型能够以有向图形式学习音素混淆模式。同时,引入一种语言特定策略,捕捉不同母语(L1)背景下的系统性发音差异。在L2-ARCTIC基准上的大量实验验证了该方法的有效性,其F1得分达到59.52%,优于多个竞争性基线。

原文摘要 · Abstract (English)

Mispronunciation Detection and Diagnosis (MDD) has gained increasing importance in computer-assisted language learning and speech technology in recent years. In this paper, we propose a method for constructing statistical graphs that enable models to learn phoneme confusion patterns represented as directed graphs. Furthermore, we introduce a language-specific strategy to capture systematic pronunciation differences across various native language (L1) backgrounds. The effectiveness of our approach is demonstrated through extensive experiments on the L2-ARCTIC benchmark, where it achieves an F1-score of 59.52%, outperforming several competitive baselines.

发音检测音素混淆多语言图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。