真实对话数据让顶尖语音识别模型表现大降,凸显现有测试集不靠谱。
ASR Benchmarking: Need for a More Representative Conversational Dataset
- 用成人电话对话构建多语言真实对话数据集
- 主流模型在对话场景下错误率显著上升
- 语音不连贯性与错误率正相关,适合真实场景研究者
自动语音识别(ASR)系统在LibriSpeech和Fleurs等常用基准上表现优异,但这些基准未能充分反映真实对话环境的复杂性——如语流不连贯、停顿、打断和多样口音。本研究基于TalkBank构建了一个多语言、非结构化的成人电话对话数据集。实验显示,多种先进ASR模型在该对话设置下性能出现显著下降。进一步发现,词错误率(Word Error Rate)与语音不连贯现象存在显著相关性,强调了开发更贴近现实对话场景的评估基准的重要性。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) systems have achieved remarkable performance on widely used benchmarks such as LibriSpeech and Fleurs. However, these benchmarks do not adequately reflect the complexities of real-world conversational environments, where speech is often unstructured and contains disfluencies such as pauses, interruptions, and diverse accents. In this study, we introduce a multilingual conversational dataset, derived from TalkBank, consisting of unstructured phone conversation between adults. Our results show a significant performance drop across various state-of-the-art ASR models when tested in conversational settings. Furthermore, we observe a correlation between Word Error Rate and the presence of speech disfluencies, highlighting the critical need for more realistic, conversational ASR benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。