arXiv:2603.14986eess.AS2026-03被引 1

通过帧间相关性估计滤波器,提升远场语音去混响的实用性能。

Deep Filter Estimation from Inter-Frame Correlations for Monaural Speech Dereverberation

  • 利用短时傅里叶变换帧间相关性,显式建模多帧深度滤波器。
  • 在REVERB数据集真实场景上SRMR指标显著提升。
  • 适合需要抗实际声学变化的语音增强应用。

远场麦克风场景下的语音去混响仍具挑战,因混响与目标信号高度相关,常导致真实环境泛化能力差。本文提出IF-CorrNet,一种基于帧间相关性的滤波器估计架构,以增强对声学变化的鲁棒性。不同于传统直接映射复谱的黑箱方法,IF-CorrNet显式利用帧间STFT相关性,为每个时频单元估计多帧深度滤波器。通过将学习目标从直接映射转向滤波器估计,网络有效约束解空间,简化训练并缓解对合成数据的过拟合。在REVERB Challenge数据集上的实验表明,IF-CorrNet在RealData上实现了显著的SRMR提升,验证了其在真实非合成环境下抑制混响与噪声的鲁棒性。

原文摘要 · Abstract (English)

Speech dereverberation in distant-microphone scenarios remains challenging due to the high correlation between reverberation and target signals, often leading to poor generalization in real-world environments. We propose IF-CorrNet, a correlation-to-filter architecture designed for robustness against acoustic variability. Unlike conventional black-box mapping methods that directly estimate complex spectra, IF-CorrNet explicitly exploits inter-frame STFT correlations to estimate multi-frame deep filters for each time-frequency bin. By shifting the learning objective from direct mapping to filter estimation, the network effectively constrains the solution space, which simplifies the training process and mitigates overfitting to synthetic data. Experimental results on the REVERB Challenge dataset demonstrate that IF-CorrNet achieves a substantial gain in the SRMR metric on RealData, confirming its robustness in suppressing reverberation and noise in practical, non-synthetic environments.

语音去混响深度学习声学鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。