轻量化语音增强模型,边缘设备实时运行效果更优
Reverse Attention for Lightweight Speech Enhancement on Edge Devices
- 在U-Net中引入软注意力门控机制提升性能
- 相比原始波形,WER降低6.24%,PESQ提升0.64分
- 适合部署在资源受限的边缘设备上
本文提出一种轻量级深度学习模型,用于在资源受限设备上实现实时语音增强。该模型采用紧凑架构,保证快速推理的同时不损失性能。核心创新在于将基于软注意力的注意力门控引入U-Net结构(该结构在分割任务中表现优异且针对GPU优化)。实验表明,该模型在语音质量与可懂度指标上表现卓越,如PESQ和词错误率(WER)等。相比未增强波形,模型实现了6.24%的WER下降和0.64分的PESQ提升,并优于同等规模的基线模型。
原文摘要 · Abstract (English)
This paper introduces a lightweight deep learning model for real-time speech enhancement, designed to operate efficiently on resource-constrained devices. The proposed model leverages a compact architecture that facilitates rapid inference without compromising performance. Key contributions include infusing soft attention-based attention gates in the U-Net architecture which is known to perform well for segmentation tasks and is optimized for GPUs. Experimental evaluations demonstrate that the model achieves competitive speech quality and intelligibility metrics, such as PESQ and Word Error Rates (WER), improving the performance of similarly sized baseline models. We are able to achieve a 6.24% WER improvement and a 0.64 PESQ score improvement over un-enhanced waveforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。