新评测指标融合语音语义,更准衡量失语者语音可懂度。
Aligning ASR Evaluation with Human and LLM Judgments: Intelligibility Metrics Using Phonetic, Semantic, and NLI Approaches
- 融合语音、语义与自然语言推理三类相似度评分
- 在可懂度数据集上与人工判断相关性达0.890
- 适合评估失语/发声障碍者的语音识别效果
传统语音识别评测指标如WER和CER难以捕捉可懂度,尤其对构音障碍和发声障碍语音,其语义一致性比字面匹配更重要。这些语音常出现音素重复、辅音不准确等错误,但听者仍能理解含义。本文指出两大挑战:(1) 现有指标无法充分反映可懂度;(2) 尽管大模型可优化识别结果,但其对失语语音的纠错效果尚未深入研究。为此,提出一种融合自然语言推理(NLI)分数、语义相似度与语音相似度的新评测方法。该方法在Speech Accessibility Project数据集上与人工判断的相关性达到0.890,显著优于传统指标,凸显以可懂度为核心评价标准的重要性。
原文摘要 · Abstract (English)
Traditional ASR metrics like WER and CER fail to capture intelligibility, especially for dysarthric and dysphonic speech, where semantic alignment matters more than exact word matches. ASR systems struggle with these speech types, often producing errors like phoneme repetitions and imprecise consonants, yet the meaning remains clear to human listeners. We identify two key challenges: (1) Existing metrics do not adequately reflect intelligibility, and (2) while LLMs can refine ASR output, their effectiveness in correcting ASR transcripts of dysarthric speech remains underexplored. To address this, we propose a novel metric integrating Natural Language Inference (NLI) scores, semantic similarity, and phonetic similarity. Our ASR evaluation metric achieves a 0.890 correlation with human judgments on Speech Accessibility Project data, surpassing traditional methods and emphasizing the need to prioritize intelligibility over error-based measures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。