ASR系统在多样语音上表现接近甚至超越人类听者。
Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results

- 对比最新ASR系统与荷兰母语者对儿童、老年人及弗拉芒口音语音的识别能力。
- 谷歌电话语音识别在多种语音中表现最优,整体性能与人类接近。
- 年龄、口音和语句长度影响性能差异,适合关注语音鲁棒性的研究者参考。
人类常被视为最佳听者,是自动语音识别(ASR)系统的上限基准。本文初步比较了先进ASR系统与荷兰母语者对多样化语音(包括荷兰儿童、老年人语音及弗拉芒口音)的识别表现。谷歌电话语音识别系统优于其他ASR系统。重要的是,这些ASR系统的表现与人类听者相当,某些情况下甚至更优。性能差异与说话人年龄、地区口音及语句长度相关。未来研究应提升ASR系统对老年化和地域口音带来的声学变化的鲁棒性。对测试样本与完整Jasmin-CGN测试集的对比表明,具体测试集选择会影响对人类与ASR性能基准的结论。
原文摘要 · Abstract (English)
Humans are often considered to be the best listeners and seen as the upper-bound performance of automatic speech recognition (ASR) systems. We present a preliminary comparison of the performances of state-of-the-art ASR systems and Dutch native listeners on the recognition of "diverse" speech, specifically Dutch child and older adults' speech and Flemish. Google Telephony outperformed the other ASR systems. Importantly, the ASR systems showed similar performance to the listeners, and in specific cases even outperformed them. Slight performance differences between the listeners and ASR systems were found related to speaker's age and regional accents and utterance length. Future research should focus on making ASR systems more robust to acoustic variability related to aging and regional accents. A comparison of ASR recognition performances on the test stimuli and the full Jasmin-CGN test sets showed the influence of the specific test sets on the conclusions regarding benchmarking human and ASR performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。