用重复使用单个模块代替堆叠,实现高效渐进式语音增强
Stack Less, Repeat More: A Block Reusing Approach for Progressive Speech Enhancement
- 通过重复使用一个处理模块替代传统堆叠,减少参数冗余
- 实验表明阶段数比独立模块数更影响性能,单模块可逐步优化语音
- 适合追求轻量化、低计算成本语音增强的场景
本文提出一种高效的语音增强方法,通过重复使用单一处理模块而非传统堆叠方式,实现渐进式优化。与增加模块数量以学习深层潜在表征不同,重复使用一个模块可在不增加参数量的情况下,逐步提升语音质量。同时,通过保持编码器和解码器浅层结构,并复用单一序列建模模块,最小化领域转换。实验表明,处理阶段数对性能的影响大于具有不同权重的模块数量。此外,所提方法在单个模块内即可实现对噪声输入的逐步精炼。进一步发现,加深编码器和解码器对学习复杂表示并无必要。结果验证了该块重用方法能实现渐进学习,为语音增强提供了一种高效替代方案。
原文摘要 · Abstract (English)
This paper presents an efficient speech enhancement (SE) approach that reuses a processing block repeatedly instead of conventional stacking. Rather than increasing the number of blocks for learning deep latent representations, repeating a single block leads to progressive refinement while reducing parameter redundancy. We also minimize domain transformation by keeping an encoder and decoder shallow and reusing a single sequence modeling block. Experimental results show that the number of processing stages is more critical to performance than the number of blocks with different weights. Also, we observed that the proposed method gradually refines a noisy input within a single block. Furthermore, with the block reuse method, we demonstrate that deepening the encoder and decoder can be redundant for learning deep complex representation. Therefore, the experimental results confirm that the proposed block reusing enables progressive learning and provides an efficient alternative for SE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。