WTFormer用小模型保留语音空间信息,提升多麦克风降噪效果
WTFormer: A Wavelet Conformer Network for MIMO Speech Enhancement with Spatial Cues Peservation
- 结合小波变换与注意力机制,捕捉多维空间特征
- 仅0.98M参数,降噪性能媲美先进系统
- 适合需要保留声源方向信息的场景
当前多通道语音增强系统多采用单输出结构,在多输入多输出(MIMO)处理中难以保持时空信号完整性。为此,本文提出一种新型神经网络WTFormer,利用小波变换的多分辨率特性与多维协同注意力机制,有效捕捉全局分布的空间特征,并采用Conformer进行时频建模。同时设计多任务损失策略并结合MUSIC算法优化训练,以最大程度保护空间信息。在LibriSpeech数据集上的实验表明,WTFormer在仅0.98M参数下,可实现与先进系统相当的去噪性能,同时更好地保留空间信息。
原文摘要 · Abstract (English)
Current multi-channel speech enhancement systems mainly adopt single-output architecture, which face significant challenges in preserving spatio-temporal signal integrity during multiple-input multiple-output (MIMO) processing. To address this limitation, we propose a novel neural network, termed WTFormer, for MIMO speech enhancement that leverages the multi-resolution characteristics of wavelet transform and multi-dimensional collaborative attention to effectively capture globally distributed spatial features, while using Conformer for time-frequency modeling. A multi task loss strategy accompanying MUSIC algorithm is further proposed for optimization training to protect spatial information to the greatest extent. Experimental results on the LibriSpeech dataset show that WTFormer can achieve comparable denoising performance to advanced systems while preserving more spatial information with only 0.98M parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。