arXiv:2507.09750cs.SDcs.LG2025-07中稿 · WASPAA25被引 3

提出频变吸声系数的合成混响数据集,提升语音增强真实感。

MB-RIRs: a Synthetic Room Impulse Response Dataset with Frequency-Dependent Absorption Coefficients

  • 基于图像源法改进,引入多频段吸声系数
  • 在真实混响上测试,音质提升0.51dB,MUSHRA得分高8.9
  • 适合语音增强、声学建模研究者使用

本文研究四种策略以提升单声道语音增强用合成混响响应(RIR)数据集的生态有效性。在传统基于图像源法(ISM)的鞋盒模型基础上,引入多频段吸声系数、声源方向性与接收器方向性;同时考虑来自SoundSpaces数据集的网格化RIR。针对每种数据集训练DeepFilternet3模型,并在真实RIR测试集上进行客观与主观评估。结果表明,采用频率依赖吸声系数的MB-RIRs在真实混响上表现更优,客观指标提升0.51dB SDR,主观评分提高8.9 MUSHRA。MB-RIRs数据集已公开免费下载。

原文摘要 · Abstract (English)

We investigate the effects of four strategies for improving the ecological validity of synthetic room impulse response (RIR) datasets for monoaural Speech Enhancement (SE). We implement three features on top of the traditional image source method-based (ISM) shoebox RIRs: multiband absorption coefficients, source directivity and receiver directivity. We additionally consider mesh-based RIRs from the SoundSpaces dataset. We then train a DeepFilternet3 model for each RIR dataset and evaluate the performance on a test set of real RIRs both objectively and subjectively. We find that RIRs which use frequency-dependent acoustic absorption coefficients (MB-RIRs) can obtain +0.51dB of SDR and a +8.9 MUSHRA score when evaluated on real RIRs. The MB-RIRs dataset is publicly available for free download.

语音增强混响生成数据集声学建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。