arXiv:2606.22022eess.AScs.SD2026-06中稿 · Interspeech 2026

用音位特征提升中文发音错误诊断准确率

Using Phonological-Level Wav2Vec2 for Mandarin Automatic Mispronunciation Detection and Diagnosis

论文配图:Using Phonological-Level Wav2Vec2 for Mandarin Automatic Mispronunciation Detection and Diagnosis
图 1 · 摘自论文原文
  • 在统一的Wav2Vec2 CTC架构中同时建模音段与声调特征
  • 误接受率降低10.1%,诊断错误率下降23.6%
  • 为二语学习者提供更细致可解释的发音反馈

自动发音错误检测与诊断(MDD)在第二语言中文发音学习中至关重要。尽管端到端(E2E)方法显著提升了音素级检测精度,但诊断反馈仍受限,因音段与声调错误未被明确区分。本文提出一种基于音位特征的MDD框架,在统一的Wav2Vec2 CTC架构中同时建模音段与声调属性。实验结果表明,相比仅基于音素的基线系统,该方法将误接受率(FAR)降低10.1%,诊断错误率(DER)降低23.6%。通过将音素分解为低层级音位成分,该方法使对二语学习者的诊断反馈更加详细且可解释。

原文摘要 · Abstract (English)

Automatic mispronunciation detection and diagnosis (MDD) plays a crucial role in L2 Mandarin pronunciation learning. While end-to-end (E2E) based MDD methods have substantially improved phoneme-level detection accuracy, diagnostic feedback remains limited, as segmental and tonal errors are not explicitly separated. In this paper, we propose a phonological feature-based MDD framework that models both segmental and tonal attributes within a unified Wav2Vec2 CTC architecture. Experimental results show that the proposed method reduces the False Acceptance Rate (FAR) by 10.1% and the Diagnostic Error Rate (DER) by 23.6% compared with the phoneme-only baseline system. By decomposing phonemes into low-level phonological components, the proposed approach enables more detailed and interpretable diagnostic feedback for L2 learners.

发音检测语音识别音位分析二语学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。