arXiv:2601.19949eess.AScs.CL2026-01被引 1

构建117.5小时混响语音数据集,每条音频附带可复现的声学参数。

RIR-Mega-Speech: A Reverberant Speech Corpus with Comprehensive Acoustic Metadata and Reproducible Evaluation

  • 用5000个模拟混响脉冲响应卷积LibriSpeech语音生成数据集。
  • 测得混响语音识别错误率比清晰语音高2.5个百分点(48%相对下降)。
  • 提供完整重建脚本和评估流程,支持结果复现。

尽管混响语音研究已持续数十年,但方法比较仍困难,因多数语料库缺乏逐文件声学标注或复现文档不足。本文提出RIR-Mega-Speech,一个约117.5小时的数据集,由LibriSpeech语音与约5000个来自RIR-Mega集合的模拟混响脉冲响应卷积生成。每条音频文件均包含通过明确、可复现方法计算出的RT60、直达-混响比(DRR)和清晰度指数(C50)。我们还提供了重建数据集和复现所有评估结果的脚本。使用Whisper small模型在1500对语音上测试,清洁语音的词错误率(WER)为5.20%(95%置信区间:4.69–5.78),混响语音为7.70%(7.04–8.35),配对增加2.50个百分点(2.06–2.98),相对下降48%。WER随RT60升高而单调上升,随DRR升高而下降,符合先前感知研究。本工作旨在为社区提供透明声学条件的标准资源,支持独立验证。仓库提供适用于Windows与Linux的一键重建指令。

原文摘要 · Abstract (English)

Despite decades of research on reverberant speech, comparing methods remains difficult because most corpora lack per-file acoustic annotations or provide limited documentation for reproduction. We present RIR-Mega-Speech, a corpus of approximately 117.5 hours created by convolving LibriSpeech utterances with roughly 5,000 simulated room impulse responses from the RIR-Mega collection. Every file includes RT60, direct-to-reverberant ratio (DRR), and clarity index ($C_{50}$) computed from the source RIR using clearly defined, reproducible procedures. We also provide scripts to rebuild the dataset and reproduce all evaluation results. Using Whisper small on 1,500 paired utterances, we measure 5.20% WER (95% CI: 4.69--5.78) on clean speech and 7.70% (7.04--8.35) on reverberant versions, corresponding to a paired increase of 2.50 percentage points (2.06--2.98). This represents a 48% relative degradation. WER increases monotonically with RT60 and decreases with DRR, consistent with prior perceptual studies. While the core finding that reverberation harms recognition is well established, we aim to provide the community with a standardized resource where acoustic conditions are transparent and results can be verified independently. The repository includes one-command rebuild instructions for both Windows and Linux environments.

语音识别混响建模数据集可复现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。