arXiv:2604.22133eess.AScs.SD2026-04

不依赖提示词的发音错误检测框架,提升识别精度与鲁棒性

Beyond Acoustic Sparsity and Linguistic Bias: A Prompt-Free Paradigm for Mispronunciation Detection and Diagnosis

论文配图:Beyond Acoustic Sparsity and Linguistic Bias: A Prompt-Free Paradigm for Mispronunciation Detection and Diagnosis
图 1 · 摘自论文原文
  • 采用帧级单调对齐的声学模型捕捉发音细微偏差
  • 在L2-ARCTIC数据集上达到71.77% F1分数
  • 适合需要高精度发音诊断的研究与教育应用

发音错误检测与诊断(MDD)需建模精细的声学差异。然而,当前基于ASR的MDD系统存在固有局限:基于CTC的模型倾向于序列级对齐,忽略瞬时发音错误线索;显式标准发音先验则使预测偏向目标发音。为此,本文提出一种无提示词框架,将声学保真度与标准引导解耦。首先引入CROTTC声学模型,强制实现单调帧级对齐以精准捕捉发音偏差;其次通过IF策略,在知识迁移原则下隐式注入发音错误信息。实验表明,CROTTC-IF在L2-ARCTIC数据集上取得71.77% F1分数,在Iqra'Eval2排行榜上达71.70% F1分数。实证分析显示,解耦声学与显式先验可显著提升MDD鲁棒性。

原文摘要 · Abstract (English)

Mispronunciation Detection and Diagnosis (MDD) requires modeling fine-grained acoustic deviations. However, current ASR-derived MDD systems often face inherent limitations. In particular, CTC-based models favor sequence-level alignments that neglect transient mispronunciation cues, while explicit canonical priors bias predictions toward intended targets. To address these bottlenecks, we propose a prompt-free framework decoupling acoustic fidelity from canonical guidance. First, we introduce CROTTC, an acoustic model enforcing monotonic, frame-level alignment to accurately capture pronunciation deviations. Second, we implicitly inject mispronunciation information via the IF strategy under the knowledge transfer principle. Experiments show CROTTC-IF achieves a 71.77% F1-score on L2-ARCTIC and 71.70% F1-score on the Iqra'Eval2 leaderboard. With empirical analysis, we demonstrate that decoupling acoustics from explicit priors provides highly robust MDD.

发音检测声学建模无提示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。