arXiv:2505.21057eess.AScs.SD2025-05中稿 · EUSIPCO 2025被引 6

轻量级变压器架构在单通道语音增强中实现顶尖性能,参数仅需原模型6%。

Study of Lightweight Transformer Architectures for Single-Channel Speech Enhancement

  • 采用时频交替的轻量化堆叠变压器结构,有效捕捉因果上下文全局依赖
  • 相比DeepFilterNet2参数减少94%,复杂度相近但性能更优
  • 结合对抗训练,适合边缘设备部署,兼顾高效与高保真语音增强

在语音增强任务中,在边缘设备的计算约束下实现最先进的性能仍面临巨大挑战。基于变换器的时序与频谱建模网络虽能提升性能,但通常带来显著的计算复杂度和模型膨胀。通过系统性消融分析,本文证明采用精简的频率-时间-频率(FTF)堆叠变换器架构可有效学习因果上下文中的全局依赖,同时避免大量计算开销。训练中引入判别器进一步提升了学习效率与增强效果,且推理阶段不增加额外复杂度。所提出的轻量级因果变换器架构结合对抗训练(LCT-GAN),在主流轻量级模型中达到最优的客观指标表现,但开销极低:相较于DeepFilterNet2仅需6%参数量,复杂度相当;对比CCFNet+(Lite)减少9%参数、10%乘加操作,性能反而更优。此外,该模型在多个常用测试集上甚至超越了更复杂的基准模型。

原文摘要 · Abstract (English)

In speech enhancement, achieving state-of-the-art (SotA) performance while adhering to the computational constraints on edge devices remains a formidable challenge. Networks integrating stacked temporal and spectral modelling effectively leverage improved architectures such as transformers; however, they inevitably incur substantial computational complexity and model expansion. Through systematic ablation analysis on transformer-based temporal and spectral modelling, we demonstrate that the architecture employing streamlined Frequency-Time-Frequency (FTF) stacked transformers efficiently learns global dependencies within causal context, while avoiding considerable computational demands. Utilising discriminators in training further improves learning efficacy and enhancement without introducing additional complexity during inference. The proposed lightweight, causal, transformer-based architecture with adversarial training (LCT-GAN) yields SoTA performance on instrumental metrics among contemporary lightweight models, but with far less overhead. Compared to DeepFilterNet2, the LCT-GAN only requires 6% of the parameters, at similar complexity and performance. Against CCFNet+(Lite), LCT-GAN saves 9% in parameters and 10% in multiply-accumulate operations yet yielding improved performance. Further, the LCT-GAN even outperforms more complex, common baseline models on widely used test datasets.

语音增强轻量模型变换器边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。