arXiv:2511.20107cs.CLcs.SD2025-11被引 2

不用训练模型,用检索法精准发现发音错误

Mispronunciation Detection and Diagnosis Without Model Training: A Retrieval-Based Approach

  • 用预训练语音识别模型做检索,不需额外训练
  • 在L2-ARCTIC上达到69.60%的F1分数
  • 适合想快速部署发音诊断系统的开发者

发音错误检测与诊断对语言学习和言语治疗至关重要。不同于传统方法需要评分模型或训练音素级模型,我们提出一种无需训练的新框架,利用预训练自动语音识别模型结合检索技术。该方法避免了音素特异性建模或额外的任务特定训练,仍能实现精确的发音错误检测与诊断。在L2-ARCTIC数据集上的实验表明,该方法取得了69.60%的优越F1分数,同时规避了模型训练的复杂性。

原文摘要 · Abstract (English)

Mispronunciation Detection and Diagnosis (MDD) is crucial for language learning and speech therapy. Unlike conventional methods that require scoring models or training phoneme-level models, we propose a novel training-free framework that leverages retrieval techniques with a pretrained Automatic Speech Recognition model. Our method avoids phoneme-specific modeling or additional task-specific training, while still achieving accurate detection and diagnosis of pronunciation errors. Experiments on the L2-ARCTIC dataset show that our method achieves a superior F1 score of 69.60% while avoiding the complexity of model training.

语音识别发音诊断零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。