用神经网络估计噪声协方差,实现无监督保向多通道语音增强。
Direction-Preserving MIMO Speech Enhancement Using a Neural Covariance Estimator

- 基于轻量级网络直接学习频域噪声协方差的归一化分解。
- 在多个下游任务中逼近理想性能,参数与计算成本显著降低。
- 适合需要保留声源方向信息的语音系统,如波束成形与双耳渲染。
多通道语音增强广泛用于麦克风阵列处理系统前端。现有方法多输出单路增强信号,而保向多输入多输出(MIMO)方法旨在生成保留空间方向特性的多通道增强信号,支持波束成形、双耳渲染和到达方向估计等下游应用。本文提出一种完全无监督的保向MIMO语音增强方法,基于神经网络对空间噪声协方差矩阵进行估计。轻量级OnlineSpatialNet估计频域噪声协方差的尺度归一化Cholesky因子,并与保向MIMO维纳滤波器结合,在增强语音的同时保持目标信号与残余噪声的空间特性。相比依赖先验信息或基于掩码的单输出系统,本方法直接实现高精度多通道协方差估计,计算复杂度低。实验表明,该方法在语音增强效果、协方差估计能力及下游任务性能上均优于基于掩码的基线,接近理想性能,且参数与计算开销大幅减少。
原文摘要 · Abstract (English)
Multichannel speech enhancement is widely used as a front-end in microphone array processing systems. While most existing approaches produce a single enhanced signal, direction-preserving multiple-input multiple-output (MIMO) methods instead aim to provide enhanced multichannel signals that retain directional properties, enabling downstream applications such as beamforming, binaural rendering, and direction-of-arrival estimation. In this work, we propose a fully blind, direction-preserving MIMO speech enhancement method based on neural estimation of the spatial noise covariance matrix. A lightweight OnlineSpatialNet estimates a scale-normalized Cholesky factor of the frequency-domain noise covariance, which is combined with a direction-preserving MIMO Wiener filter to enhance speech while preserving the spatial characteristics of both target and residual noise. In contrast to prior approaches relying on oracle information or mask-based covariance estimation for single-output systems, the proposed method directly targets accurate multichannel covariance estimation with low computational complexity. Experimental results show improved speech enhancement, covariance estimation capability, and performance in downstream tasks over a mask-based baseline, approaching oracle performance with significantly fewer parameters and computational cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。