arXiv:2510.18917eess.ASeess.SP2025-10被引 2

构建大规模仿真混响数据集,助力语音降噪与声学建模研究

RIR-Mega: a large-scale simulated room impulse response dataset for machine learning and room acoustics modeling

  • 基于紧凑元数据格式生成5万条仿真混响响应,支持机器高效处理
  • 用轻量特征训练随机森林,预测混响时间误差仅0.013秒(MAE)
  • 提供可直接调用的Hugging Face数据加载器,适合语音与声学研究者

混响冲击响应是语音去混响、鲁棒语音识别、声源定位和房间声学建模的核心资源。我们提出RIR-Mega,一个大规模仿真混响响应数据集,采用简洁、机器友好的元数据格式,并配备验证与复用工具。数据集附带Hugging Face Datasets加载器、元数据校验脚本与哈希校验工具,以及一个参考回归基线模型,可从波形中预测RT60目标值。在36,000条训练集和4,000条验证集上,使用轻量时间与频域特征的简单随机森林模型,达到平均绝对误差约0.013秒、均方根误差约0.022秒。我们已在Hugging Face托管包含1,000条线性阵列和3,000条环形阵列的子集,用于流式传输与快速测试,并将完整的50,000条混响响应档案永久保存于Zenodo。数据集与代码均公开,支持可复现研究。

原文摘要 · Abstract (English)

Room impulse responses are a core resource for dereverberation, robust speech recognition, source localization, and room acoustics estimation. We present RIR-Mega, a large collection of simulated RIRs described by a compact, machine friendly metadata schema and distributed with simple tools for validation and reuse. The dataset ships with a Hugging Face Datasets loader, scripts for metadata checks and checksums, and a reference regression baseline that predicts RT60 like targets from waveforms. On a train and validation split of 36,000 and 4,000 examples, a small Random Forest on lightweight time and spectral features reaches a mean absolute error near 0.013 s and a root mean square error near 0.022 s. We host a subset with 1,000 linear array RIRs and 3,000 circular array RIRs on Hugging Face for streaming and quick tests, and preserve the complete 50,000 RIR archive on Zenodo. The dataset and code are public to support reproducible studies.

语音处理混响建模数据集机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。