构建1.5小时荷兰语真实场景语音数据集,用于评估语音识别与增强模型性能。
A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition
- 采集80人于嘈杂公共室内环境的半自发语音,使用四麦克风阵列
- 8个主流语音识别模型在该数据集上词错误率低于22%(5/8)
- 发现现代单通道语音增强未提升识别效果,凸显真实场景评估重要性
我们提出DRES:一个包含1.5小时荷兰语真实场景半自发语音的数据集,涵盖80名说话者在嘈杂公共室内环境中的录音。该数据集旨在评估先进自动语音识别(ASR)和语音增强(SE)模型在现实场景中的表现——即人在有背景人声与噪声的公共场所讲话。语音通过四通道线性麦克风阵列采集。本研究评估了五种知名单通道语音增强算法的语音质量,以及八种主流离线ASR模型在应用SE前后的识别性能。结果显示,在严苛条件下,8个模型中有5个词错误率(WER)低于22%。与近期研究不同,我们未观察到现代单通道语音增强对ASR性能的正向影响,强调了在真实环境中评估的重要性。
原文摘要 · Abstract (English)
We present DRES: a 1.5-hour Dutch realistic elicited (semi-spontaneous) speech dataset from 80 speakers recorded in noisy, public indoor environments. DRES was designed as a test set for the evaluation of state-of-the-art (SotA) automatic speech recognition (ASR) and speech enhancement (SE) models in a real-world scenario: a person speaking in a public indoor space with background talkers and noise. The speech was recorded with a four-channel linear microphone array. In this work we evaluate the speech quality of five well-known single-channel SE algorithms and the recognition performance of eight SotA off-the-shelf ASR models before and after applying SE on the speech of DRES. We found that five out of the eight ASR models have WERs lower than 22\% on DRES, despite the challenging conditions. In contrast to recent work, we did not find a positive effect of modern single-channel SE on ASR performance, emphasizing the importance of evaluating in realistic conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。