arXiv:2506.12067eess.AScs.AI2025-06中稿 · Interspeech 2025被引 4

用对数几率替代概率,提升发音错误检测准确率

Evaluating Logit-Based GOP Scores for Mispronunciation Detection

  • 用对数几率替代后验概率计算发音评分
  • 最大对数几率评分与人工评估相关性最强
  • 混合方法结合不确定性建模,适合多语种发音评估

发音评估依赖于发音质量得分(GOP),传统上基于Softmax后验概率计算。然而,后验概率存在过度自信和音素区分度差的问题,限制了其效果。本研究比较了基于对数几率的GOP与基于概率的GOP在发音错误检测中的表现。实验基于荷蘭語和中文母語者說英語的兩個L2英語語料庫,評估分類性能與人工評分的相關性。結果顯示,對數几率方法在分類任務中優於概率方法,但效果依賴數據集特徵。最大對數几率GOP與人為感知最一致,而融合不同GOP得分的方法則平衡了概率與對數几率特性。研究表明,結合不確定性建模與音素特異加權的混合型GOP方法能有效提升發音評估性能。

原文摘要 · Abstract (English)

Pronunciation assessment relies on goodness of pronunciation (GOP) scores, traditionally derived from softmax-based posterior probabilities. However, posterior probabilities may suffer from overconfidence and poor phoneme separation, limiting their effectiveness. This study compares logit-based GOP scores with probability-based GOP scores for mispronunciation detection. We conducted our experiment on two L2 English speech datasets spoken by Dutch and Mandarin speakers, assessing classification performance and correlation with human ratings. Logit-based methods outperform probability-based GOP in classification, but their effectiveness depends on dataset characteristics. The maximum logit GOP shows the strongest alignment with human perception, while a combination of different GOP scores balances probability and logit features. The findings suggest that hybrid GOP methods incorporating uncertainty modeling and phoneme-specific weighting improve pronunciation assessment.

发音评估对数几率语音识别机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。