arXiv:2608.25384eess.AS2026-08中稿 · publication at ISC…

基于音系特征的语音纠错框架,提升中文二语学习者发音诊断精度。

Mandarin Humorous Homophone Recognition and Disambiguation in Automatic Speech Recognition

论文配图:Mandarin Humorous Homophone Recognition and Disambiguation in Automatic Speech Recognition
图 1 · 摘自论文原文
  • 在Wav2Vec2-CTC架构中联合建模音段与声调特征
  • 误判率降低10.1%,诊断错误率下降23.6%
  • 可为二语学习者提供更清晰可解释的发音反馈

自动发音错误检测与诊断(MDD)在第二语言汉语发音学习中至关重要。尽管端到端(E2E)方法显著提升了音素级检测准确率,但诊断反馈仍受限于音段与声调错误未被明确区分。本文提出一种基于音系特征的MDD框架,在统一的Wav2Vec2-CTC架构中同时建模音段与声调属性。实验表明,相比仅使用音素的基线系统,该方法将误判率(FAR)降低10.1%,诊断错误率(DER)降低23.6%。通过将音素分解为低层音系成分,该方法使对二语学习者的诊断反馈更加细致且可解释。

原文摘要 · Abstract (English)

Automatic mispronunciation detection and diagnosis (MDD) plays a crucial role in L2 Mandarin pronunciation learning. While end-to-end (E2E) based MDD methods have substantially improved phoneme-level detection accuracy, diagnostic feedback remains limited, as segmental and tonal errors are not explicitly separated. In this paper, we propose a phonological feature-based MDD framework that models both segmental and tonal attributes within a unified Wav2Vec2-CTC architecture. Experimental results show that the proposed method reduces the False Acceptance Rate (FAR) by 10.1% and the Diagnostic Error Rate (DER) by 23.6% compared with the phoneme-only baseline system. By decomposing phonemes into low-level phonological components, the proposed approach enables more detailed and interpretable diagnostic feedback for L2 learners.

语音识别发音诊断音系特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。