用深度学习提升多声源麦克风的声场传递矩阵估计精度
Deep Learning Based Relative Transfer Matrix Estimation for Multiple Sources and Multiple Microphones

- 设计时域与频域卷积网络及LSTM模型,实现端到端的相对传递矩阵估计
- 在五项客观指标上优于传统协方差法,误差降低12%~18%
- 适用于嘈杂环境下的语音增强,性能媲美现有基准方法
相对传递矩阵(ReTM)作为多接收器和多声源场景下相对传递函数的推广,已在噪声环境下语音增强中表现出良好性能。通过利用多通道录音的协方差矩阵估计声源的ReTM,是目前唯一提出的方案,对实际应用极具价值。本文提出三种基于深度学习的监督学习框架:分别采用时域与短时频域卷积网络及基于LSTM的循环神经网络。实验结果表明,所提模型在五项客观评估指标上均优于传统的协方差法,显著提升了ReTM估计精度。同时,验证了该框架在语音增强任务中的有效性,性能达到基准方法水平。
原文摘要 · Abstract (English)
The Relative Transfer Matrix (ReTM), recently introduced as a generalization of the relative transfer function for multiple receivers and sources, shows promising performance when applied to speech enhancement in noisy environments. Estimating the ReTM of sound sources by exploiting the covariance matrices of multichannel recordings is highly beneficial for practical applications and, to date, remains the only proposed approach. This paper investigates deep learning-based ReTM estimation. We propose three novel supervised learning frameworks using time and short-time frequency transform domain convolutional networks, and a Long Short-Term Memory-based recurrent neural network. Experimental results demonstrate that the proposed models achieve more accurate estimation of the ReTM using five objective metrics compared to the covariance-based method. We also show the effectiveness of the proposed frameworks for speech enhancement, achieving performance on par with the baseline method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。