arXiv:2601.13140eess.AS2026-01

用多麦克风注意力机制提升语音降噪的生成效果

AMDM-SE: Attention-based Multichannel Diffusion Model for Speech Enhancement

  • 引入跨通道时频注意力块,捕捉多麦克风间的空间信息
  • 在CHiME-3上优于单通道与无注意力的多通道基线模型
  • 适合研究语音增强与扩散模型融合的学者参考

扩散模型在从噪声输入重建图像方面取得显著进展,类似思想已被应用于将时频表示视为图像的语音增强。随着多麦克风设备的普及,本文将先进的基于扩散的方法扩展至多通道输入以提升性能。目前基于扩散的多通道语音增强仍处于初期阶段,先前工作对注意力等先进机制的空间建模利用有限,本文正是针对这一空白提出AMD-SE:一种面向语音降噪的注意力型多通道扩散模型。该模型通过新颖的跨通道时频注意力模块,有效利用通道间空间信息,在生成式扩散框架中实现精细信号细节的忠实重建。在CHiME-3基准测试中,AMDM-SE超越单通道扩散基线、无注意力的多通道模型以及强基线的DNN预测方法。模拟数据实验进一步验证了所提多通道注意力机制的重要性。结果表明,将针对性的多通道注意力融入扩散模型可显著提升降噪效果。尽管基于扩散的多通道语音增强仍是新兴领域,本工作为该方向提供了新的互补性方法。

原文摘要 · Abstract (English)

Diffusion models have recently achieved impressive results in reconstructing images from noisy inputs, and similar ideas have been applied to speech enhancement by treating time-frequency representations as images. With the ubiquity of multi-microphone devices, we extend state-of-the-art diffusion-based methods to exploit multichannel inputs for improved performance. Multichannel diffusion-based enhancement remains in its infancy, with prior work making limited use of advanced mechanisms such as attention for spatial modeling - a gap addressed in this paper. We propose AMDM-SE, an Attention-based Multichannel Diffusion Model for Speech Enhancement, designed specifically for noise reduction. AMDM-SE leverages spatial inter-channel information through a novel cross-channel time-frequency attention block, enabling faithful reconstruction of fine-grained signal details within a generative diffusion framework. On the CHiME-3 benchmark, AMDM-SE outperforms both a single-channel diffusion baseline and a multichannel model without attention, as well as a strong DNN-based predictive method. Simulated-data experiments further underscore the importance of the proposed multichannel attention mechanism. Overall, our results show that incorporating targeted multichannel attention into diffusion models substantially improves noise reduction. While multichannel diffusion-based speech enhancement is still an emerging field, our work contributes a new and complementary approach to the growing body of research in this direction.

语音增强扩散模型多麦克风注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。