轻量级语音增强模型,高效融合多尺度声学特征。
EffiFusion-GAN: Efficient Fusion Generative Adversarial Network for Speech Enhancement
- 用深度可分离卷积在多尺度块中捕捉声学特征
- 双归一化注意力+残差优化,提升训练稳定性和收敛速度
- 动态剪枝压缩模型,适合低资源设备部署
我们提出EffiFusion-GAN(高效融合生成对抗网络),一种轻量但强大的语音增强模型。该模型在多尺度模块中引入深度可分离卷积,以高效捕捉多样化的声学特征。通过双重归一化和残差精炼的增强注意力机制,进一步提升训练稳定性与收敛性能。此外,采用动态剪枝策略,在保持性能的同时减少模型规模,使系统适用于资源受限环境。在公开数据集VoiceBank+DEMAND上的实验表明,EffiFusion-GAN在相同参数量下取得3.45的PESQ得分,优于现有模型。
原文摘要 · Abstract (English)
We introduce EffiFusion-GAN (Efficient Fusion Generative Adversarial Network), a lightweight yet powerful model for speech enhancement. The model integrates depthwise separable convolutions within a multi-scale block to capture diverse acoustic features efficiently. An enhanced attention mechanism with dual normalization and residual refinement further improves training stability and convergence. Additionally, dynamic pruning is applied to reduce model size while maintaining performance, making the framework suitable for resource-constrained environments. Experimental evaluation on the public VoiceBank+DEMAND dataset shows that EffiFusion-GAN achieves a PESQ score of 3.45, outperforming existing models under the same parameter settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。