arXiv:2409.06456cs.SDeess.AS2024-09被引 6

用注意力机制提升语音增强中空间协方差矩阵估计精度。

Attention-Based Beamformer For Multi-Channel Speech Enhancement

  • 引入注意力机制动态估计语音与噪声的空间协方差矩阵。
  • 在多种条件下均优于基线方法,计算量和参数更少。
  • 适合需要实时部署的多通道语音增强场景。

最小方差无失真响应(MVDR)是一种经典的自适应波束成形器,理论上可保证目标方向信号无失真传输,因此广泛应用于实际场景。其降噪性能取决于噪声与语音空间协方差矩阵(SCMs)估计的准确性。通常使用时频掩码计算这些矩阵。然而,多数基于掩码的波束成形方法假设源信号静止,忽略移动源情况,导致性能下降。本文提出一种基于注意力机制的方法,用于计算语音与噪声的SCMs,再应用MVDR获得增强语音。为充分融合空间信息,采用原位卷积算子和频率无关的LSTM辅助SCMs估计,模型以端到端方式优化。实验表明,该方法在多种条件下均优于基线,且计算开销和参数量更少。

原文摘要 · Abstract (English)

Minimum Variance Distortionless Response (MVDR) is a classical adaptive beamformer that theoretically ensures the distortionless transmission of signals in the target direction, which makes it popular in real applications. Its noise reduction performance actually depends on the accuracy of the noise and speech spatial covariance matrices (SCMs) estimation. Time-frequency masks are often used to compute these SCMs. However, most mask-based beamforming methods typically assume that the sources are stationary, ignoring the case of moving sources, which leads to performance degradation. In this paper, we propose an attention-based mechanism to calculate the speech and noise SCMs and then apply MVDR to obtain the enhanced speech. To fully incorporate spatial information, the inplace convolution operator and frequency-independent LSTM are applied to facilitate SCMs estimation. The model is optimized in an end-to-end manner. Experiments demonstrate that the proposed method outperforms baselines with reduced computation and fewer parameters under various conditions.

语音增强波束成形注意力机制多通道

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。