用深度网络分步去噪去混响,重建语音频谱以识别房间脉冲响应。
Blind Room Impulse Response Identification via Reverberant Speech Spectrum Reconstruction
- 分步处理:先去噪再去混响,通过频谱重建估计声学传输函数。
- 在真实录音上实现当前最优的盲态房间脉冲响应识别性能。
- 适合语音增强、声学定位等需要精确声学建模的场景。
本文提出Rec-RIR方法,用于盲态房间脉冲响应(RIR)识别。基于卷积传递函数(CTF)近似,设计了一个多任务深度神经网络,依次从语音录音中去除噪声和混响,并通过重构混响语音频谱来估计CTF滤波器。随后,采用伪侵入式测量过程,模拟常见侵入式RIR测量流程,将CTF滤波器转换为RIR。实验结果表明,Rec-RIR在盲态RIR识别任务上达到当前最优(SOTA)性能。
原文摘要 · Abstract (English)
This paper proposes Rec-RIR for blind room impulse response (RIR) identification. Based on the convolutive transfer function (CTF) approximation, we propose a multi-task deep neural network, which sequentially removes noise and reverberation from speech recording, and estimates the CTF filter by reverberant speech spectrum reconstruction. Subsequently, a pseudo intrusive measurement process is employed to convert the CTF filter into RIR by simulating a common intrusive RIR measurement procedure. Experimental results demonstrate that Rec-RIR achieves state-of-the-art (SOTA) performance in blind RIR identification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。