用深度网络联合处理空间与频谱信息,有效抑制麦克风阵列中的空间混叠问题。
An Analysis of Joint Nonlinear Spatial Filtering for Spatial Aliasing Reduction
- 采用非线性深度网络联合进行空间与频谱处理
- 在大间距麦克风阵列下显著降低空间混叠现象
- 适合大尺寸阵列、高频率场景的语音增强应用
传统线性空间滤波器在语音增强中的性能受限于麦克风阵列的物理尺寸和通道数。例如,在麦克风间距较大且频率较高时,会出现空间混叠,导致非目标方向信号被错误增强。近期研究提出用非线性深度神经网络替代线性波束成形器,实现联合空间-频谱处理。本文表明,此类方法不仅能提升仪器评估指标,更在抑制空间混叠方面表现优异。特别地,联合空间与时频处理比仅做空间处理或分离处理更具鲁棒性。结果进一步支持在多通道语音增强中使用深度非线性网络,尤其在麦克风间距较大的情况下,除能应对非高斯噪声和多说话人外,还能有效缓解空间混叠问题。
原文摘要 · Abstract (English)
The performance of traditional linear spatial filters for speech enhancement is constrained by the physical size and number of channels of microphone arrays. For instance, for large microphone distances and high frequencies, spatial aliasing may occur, leading to unwanted enhancement of signals from non-target directions. Recently, it has been proposed to replace linear beamformers by nonlinear deep neural networks for joint spatial-spectral processing. While it has been shown that such approaches result in higher performance in terms of instrumental quality metrics, in this work we highlight their ability to efficiently handle spatial aliasing. In particular, we show that joint spatial and tempo-spectral processing is more robust to spatial aliasing than traditional approaches that perform spatial processing alone or separately with tempo-spectral filtering. The results provide another strong motivation for using deep nonlinear networks in multichannel speech enhancement, beyond their known benefits in managing non-Gaussian noise and multiple speakers, especially when microphone arrays with rather large microphone distances are used.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。