用神经网络联合优化时频域波束成形权重,提升多人混响中目标语音提取效果。
Neural Network-Based Time-Frequency-Bin-Wise Linear Combination of Beamformers for Underdetermined Target Source Extraction
- 通过神经网络预测时频相干的加权系数,避免独立处理每个时频点
- 在双麦克风多干扰场景下性能超越传统方法,接近需噪声先验的最优解
- 无需显式估计噪声协方差,适用于实际部署中噪声未知场景
从欠定混合信号中提取目标源对波束成形方法构成挑战。近期提出的时频域切换(TFS)与线性组合(TFLC)策略通过在每个时频(TF)单元组合多个波束成形器并最小化输出功率来缓解问题,但对每个时频点独立决策会削弱时频相干性,导致不连续性从而降低提取性能。本文提出一种新型神经网络驱动的时频域线性组合(NN-TFLC)框架,构建无需显式噪声协方差估计的最小功率无失真响应(MPDR)波束成形器。网络编码混合信号与波束成形输出,并通过交叉注意力机制预测具有时频一致性的加权系数。在双麦克风含多干扰源的混合场景下,NN-TFLC-MPDR始终优于TFS/TFLC-MPDR,且性能媲美依赖噪声先验的最小方差无失真响应(MVDR)基线方法。
原文摘要 · Abstract (English)
Extracting a target source from underdetermined mixtures is challenging for beamforming approaches. Recently proposed time-frequency-bin-wise switching (TFS) and linear combination (TFLC) strategies mitigate this by combining multiple beamformers in each time-frequency (TF) bin and choosing combination weights that minimize the output power. However, making this decision independently for each TF bin can weaken temporal-spectral coherence, causing discontinuities and consequently degrading extraction performance. In this paper, we propose a novel neural network-based time-frequency-bin-wise linear combination (NN-TFLC) framework that constructs minimum power distortionless response (MPDR) beamformers without explicit noise covariance estimation. The network encodes the mixture and beamformer outputs, and predicts temporally and spectrally coherent linear combination weights via a cross-attention mechanism. On dual-microphone mixtures with multiple interferers, NN-TFLC-MPDR consistently outperforms TFS/TFLC-MPDR and achieves competitive performance with TFS/TFLC built on the minimum variance distortionless response (MVDR) beamformers that require noise priors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。