用双路径下采样上采样提升单声道语音增强效率与效果
ZipEnhancer: Dual-Path Down-Up Sampling-based Zipformer for Monaural Speech Enhancement
- 采用时域与频域双路径下采样上采样结构,降低计算开销
- 在DNS 2020和Voicebank+DEMAND数据集上达到PESQ 3.69和3.63新高
- 仅用204万参数和624.1亿浮点运算,适合资源受限场景
与传统三轴隐藏层特征建模不同,双路径时域与频域语音增强模型虽参数少但计算量大,因其隐藏层特征具四轴特性。本文提出ZipEnhancer——基于双路径下采样上采样的Zipformer,用于单声道语音增强。引入核心模块ZipformerBlock及对称式下采样堆栈设计,实现高效特征压缩与恢复。同时提出ScaleAdam优化器与Eden学习率调度器以进一步提升性能。在DNS 2020 Challenge和Voicebank+DEMAND数据集上取得新基准表现,PESQ分别为3.69与3.63,仅使用204万参数与624.1亿次浮点运算,优于同复杂度方法。
原文摘要 · Abstract (English)
In contrast to other sequence tasks modeling hidden layer features with three axes, Dual-Path time and time-frequency domain speech enhancement models are effective and have low parameters but are computationally demanding due to their hidden layer features with four axes. We propose ZipEnhancer, which is Dual-Path Down-Up Sampling-based Zipformer for Monaural Speech Enhancement, incorporating time and frequency domain Down-Up sampling to reduce computational costs. We introduce the ZipformerBlock as the core block and propose the design of the Dual-Path DownSampleStacks that symmetrically scale down and scale up. Also, we introduce the ScaleAdam optimizer and Eden learning rate scheduler to improve the performance further. Our model achieves new state-of-the-art results on the DNS 2020 Challenge and Voicebank+DEMAND datasets, with a perceptual evaluation of speech quality (PESQ) of 3.69 and 3.63, using 2.04M parameters and 62.41G FLOPS, outperforming other methods with similar complexity levels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。