arXiv:2606.24137eess.AScs.SD2026-06中稿 · INTERSPEECH 2026

用神经网络动态调整噪声增益,提升语音增强鲁棒性

Joint Learning of Covariance Estimation and White Noise Gain for Robust MVDR Beamforming

论文配图:Joint Learning of Covariance Estimation and White Noise Gain for Robust MVDR Beamforming
图 1 · 摘自论文原文
  • 用深度网络联合估计噪声协方差和频率相关噪声增益阈值
  • 在多种噪声条件下,语音质量与可懂度显著优于固定阈值方法
  • 适合需要自适应抗噪的实时语音处理系统

最小方差无失真响应(MVDR)波束成形器因强噪声抑制能力被广泛用于多通道语音增强,但其性能易受麦克风自噪声和阵列失配影响。现有方法通常依赖固定手动调参的白噪声增益(WNG)阈值或对角加载,难以应对未知或时变声学环境。本文提出一种数据驱动的MVDR框架,通过深度神经网络自适应估计WNG约束。网络联合预测时频域噪声掩码以实现协方差估计,并输出频率相关的WNG阈值,从而实现动态鲁棒性-方向性调控。框架中集成可微分的鲁棒MVDR层,支持端到端优化。实验表明,在不同噪声条件下,该方法在语音质量与可懂度上均持续优于传统固定WNG的MVDR方法。

原文摘要 · Abstract (English)

The minimum variance distortionless response (MVDR) beamformer is widely used for multichannel speech enhancement due to strong noise suppression while preserving target signals. In practice, its performance is sensitive to microphone self-noise and array mismatches. Existing approaches typically rely on fixed, manually tuned WNG thresholds or diagonal loading, leading to suboptimal performance under unknown or time-varying acoustic conditions. This paper proposes a data-driven MVDR framework that adaptively estimates the WNG constraint using a deep neural network. The network jointly predicts a time-frequency noise mask for covariance estimation and a frequency-dependent WNG threshold, enabling dynamic robustness-directivity control. A differentiable robust MVDR layer is integrated into the framework, allowing end-to-end optimization. Experiments demonstrate consistent improvements in speech quality and intelligibility over conventional fixed-WNG MVDR methods.

语音增强波束成形深度学习自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。