arXiv:2502.19906eess.AScs.SD2025-02被引 16

用素数核卷积实现高效多尺度频谱学习,显著提升单通道语音增强效果。

PrimeK-Net: Multi-scale Spectral Learning via Group Prime-Kernel Convolutional Neural Networks for Single Channel Speech Enhancement

  • 引入分组素数核卷积,实现低复杂度的多粒度频谱建模。
  • 在VoiceBank+Demand数据集上达到3.61的PESQ得分,仅需141万参数。
  • 适合关注语音增强效率与性能平衡的研究者与工程师。

单通道语音增强是一项具有挑战性的病态问题,旨在从退化信号中恢复清晰语音。现有研究已证明卷积神经网络(CNN)与Transformer结合在该任务中表现优异,但现有框架在计算效率方面仍有不足,且未充分考虑频谱的天然多尺度特性,也未充分发挥CNN的潜力。为此,本文提出深度可分离膨胀密集块(DSDDB)和分组素数核前馈通道注意力(GPFCA)模块。DSDDB提升了编码器/解码器的参数与计算效率;GPFCA模块替代Conformer结构,以线性复杂度提取深层时频特征。其基于提出的分组素数核前馈网络(GPFN),融合长、中、短程感受野,并利用素数性质避免周期性重叠效应。实验表明,所提PrimeK-Net在VoiceBank+Demand数据集上达到当前最优性能,PESQ得分为3.61,仅需141万参数。

原文摘要 · Abstract (English)

Single-channel speech enhancement is a challenging ill-posed problem focused on estimating clean speech from degraded signals. Existing studies have demonstrated the competitive performance of combining convolutional neural networks (CNNs) with Transformers in speech enhancement tasks. However, existing frameworks have not sufficiently addressed computational efficiency and have overlooked the natural multi-scale distribution of the spectrum. Additionally, the potential of CNNs in speech enhancement has yet to be fully realized. To address these issues, this study proposes a Deep Separable Dilated Dense Block (DSDDB) and a Group Prime Kernel Feedforward Channel Attention (GPFCA) module. Specifically, the DSDDB introduces higher parameter and computational efficiency to the Encoder/Decoder of existing frameworks. The GPFCA module replaces the position of the Conformer, extracting deep temporal and frequency features of the spectrum with linear complexity. The GPFCA leverages the proposed Group Prime Kernel Feedforward Network (GPFN) to integrate multi-granularity long-range, medium-range, and short-range receptive fields, while utilizing the properties of prime numbers to avoid periodic overlap effects. Experimental results demonstrate that PrimeK-Net, proposed in this study, achieves state-of-the-art (SOTA) performance on the VoiceBank+Demand dataset, reaching a PESQ score of 3.61 with only 1.41M parameters.

语音增强多尺度学习素数核轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。