arXiv:2502.11462eess.AScs.LG2025-02中稿 · ICASSP 2025被引 4

轻量级语音增强网络,高效捕捉频带间信息。

LMFCA-Net: A Lightweight Model for Multi-Channel Speech Enhancement with Efficient Narrow-Band and Cross-Band Attention

  • 分离式全连接注意力机制,无需循环单元
  • 计算量和延迟显著降低,性能接近顶尖方法
  • 适合终端设备部署,兼顾效率与效果

基于深度学习的端到端多通道语音增强方法通过利用子带、跨带和空间信息取得了优异性能。然而,这些方法通常需要大量计算资源,限制了其在终端设备上的实际应用。本文提出一种轻量级多通道语音增强网络(LMFCA-Net),引入时轴解耦全连接注意力(T-FCA)和频轴解耦全连接注意力(F-FCA)机制,有效捕获长程窄带和跨带信息,且无需循环单元。实验结果表明,LMFCA-Net在性能上可与当前最优方法相媲美,同时显著降低计算复杂度和延迟,是一种极具潜力的实际应用解决方案。

原文摘要 · Abstract (English)

Deep learning based end-to-end multi-channel speech enhancement methods have achieved impressive performance by leveraging sub-band, cross-band, and spatial information. However, these methods often demand substantial computational resources, limiting their practicality on terminal devices. This paper presents a lightweight multi-channel speech enhancement network with decoupled fully connected attention (LMFCA-Net). The proposed LMFCA-Net introduces time-axis decoupled fully-connected attention (T-FCA) and frequency-axis decoupled fully-connected attention (F-FCA) mechanisms to effectively capture long-range narrow-band and cross-band information without recurrent units. Experimental results show that LMFCA-Net performs comparably to state-of-the-art methods while significantly reducing computational complexity and latency, making it a promising solution for practical applications.

语音增强轻量模型注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。