研究现代语音识别模型与人耳听觉效果的匹配度,发现复杂模型更接近真实感知。
Too Good to Be True: A Study on Modern Automatic Speech Recognition for the Evaluation of Speech Enhancement

- 用大噪声数据训练的现代ASR模型更贴近人类听觉表现
- 基于转换器的模型生成转录最可靠,相关性最高
- 但其抗噪能力可能掩盖语音增强的真实效果,适合评估者参考
语音增强(SE)系统通常通过多种仪器指标进行评估,其中自动语音识别(ASR)系统常以词错误率(WER)作为评价标准。然而,WER结果高度依赖于ASR系统和文本归一化流程的选择。本文研究现代ASR模型与人类对增强语音识别能力之间的相关性。听觉实验表明,经过大规模噪声数据训练并集成语言模型的现代ASR模型,相较于简单模型,与人类WER的相关性更高,其中转换器(transducer)模型提供最可靠的转录结果。然而,我们同时发现这些模型对噪声的鲁棒性以及上下文使用能力,在聚焦声学性能的增强评估中可能产生误导性信息。
原文摘要 · Abstract (English)
Speech enhancement (SE) systems are typically evaluated using a variety of instrumental metrics. The use of automatic speech recognition (ASR) systems to evaluate SE performance is common in literature, usually in terms of word error rate (WER). However, WER scores depend heavily on the choice of ASR system and text normalization pipeline. In this paper, we investigate how modern ASR models correlate with human recognition of enhanced speech. A listening experiment reveals that modern ASR models with large-scale noisy training and embedded language models correlate more with human WER than simpler ones, with a transducer model providing the most reliable transcriptions. Nevertheless, we also show that these models' robustness to noise and use of context can be uninformative to an acoustics-focused evaluation of enhancement performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。