构建真实混响语音数据集,评估语音识别模型在不同声学环境下的鲁棒性。
Whisper-RIR-Mega: A Paired Clean-Reverberant Speech Benchmark for ASR Robustness to Room Acoustics
- 用真实房间脉冲响应混合干净语音,生成成对的清晰与混响语音数据。
- 大模型(Whisper-large-v3)混响误差增幅最小,小模型(Whisper-tiny)最大达15.50个百分点。
- 适合研究语音识别鲁棒性、声学建模或语音增强的科研人员使用。
我们提出 Whisper-RIR-Mega,一个用于评估自动语音识别(ASR)对房间声学鲁棒性的配对清晰-混响语音基准数据集。每个样本将一个 LibriSpeech 语音片段与其经过 RIR-Mega 数据集中真实房间脉冲响应卷积后的版本配对,并按混响时间(RT60)和直达声与混响声比(DRR)进行分层划分。我们在 1600 个测试样本上评估了五个 Whisper 模型(tiny 到 large-v3),报告了在清晰与混响条件下单词错误率(WER)和字符错误率(CER)。所有模型在混响下性能均下降;混响带来的 WER 增幅为 2.31 至 15.50 个百分点,取决于模型大小。Whisper-large-v3 的增益最小,Whisper-tiny 最大。我们公开数据集、评估代码和基线结果,以支持可复现的鲁棒语音识别研究。
原文摘要 · Abstract (English)
We introduce Whisper-RIR-Mega, a benchmark dataset of paired clean and reverberant speech for evaluating automatic speech recognition (ASR) robustness to room acoustics. Each sample pairs a clean LibriSpeech utterance with the same utterance convolved with a real room impulse response from the RIR-Mega corpus, with stratified splits by reverberation time (RT60) and direct-to-reverberant ratio (DRR). We evaluate five Whisper models (tiny through large-v3) on 1600 test samples and report word error rate (WER) and character error rate (CER) under clean and reverberant conditions. Reverberation consistently degrades performance across all model sizes; the reverb penalty in WER ranges from 2.31 to 15.50 percentage points depending on the model. Whisper-large-v3 shows the smallest penalty; Whisper-tiny shows the largest. We release the dataset, evaluation code, and baseline results to support reproducible research on robust ASR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。