arXiv:2509.23832eess.AScs.SD2025-09

轻量级语音增强模型LORT,兼顾效果与效率。

LORT: Locally Refined Convolution and Taylor Transformer for Monaural Speech Enhancement

  • 融合局部精修卷积与泰勒注意力,提升特征建模能力。
  • 仅0.96M参数,在VCTK+DEMAND和DNS数据集上达顶尖性能。
  • 适合资源受限场景,如移动端或嵌入式语音处理。

在保持低参数量和计算复杂度的前提下实现卓越的语音增强性能,仍是语音增强领域的挑战。本文提出LORT,一种结合空间-通道增强泰勒变压器与局部精修卷积的新型架构,用于高效且鲁棒的单声道语音增强。我们设计了增强空间-通道注意力(SCEA)的泰勒多头自注意力(T-MSA)模块,促进通道间信息交换,缓解基于泰勒的变压器固有的空间注意力局限性。为补充全局建模,进一步提出局部精修卷积(LRC)模块,集成卷积前馈层、时频密集局部卷积和门控单元,以捕捉细粒度局部细节。LORT基于类似U-Net的编码器-解码器结构,编码器仅含16个输出通道,通过交替下采样和上采样操作处理噪声输入,使用多分辨率T-MSA模块进行建模。增强后的幅度谱和相位谱独立解码,并通过联合考虑幅度、复数、相位、判别器和一致性目标的复合损失函数进行优化。在VCTK+DEMAND和DNS Challenge数据集上的实验结果表明,LORT仅用0.96M参数即达到或超越现有最先进模型性能,验证了其在计算资源受限的真实场景中的有效性。

原文摘要 · Abstract (English)

Achieving superior enhancement performance while maintaining a low parameter count and computational complexity remains a challenge in the field of speech enhancement. In this paper, we introduce LORT, a novel architecture that integrates spatial-channel enhanced Taylor Transformer and locally refined convolution for efficient and robust speech enhancement. We propose a Taylor multi-head self-attention (T-MSA) module enhanced with spatial-channel enhancement attention (SCEA), designed to facilitate inter-channel information exchange and alleviate the spatial attention limitations inherent in Taylor-based Transformers. To complement global modeling, we further present a locally refined convolution (LRC) block that integrates convolutional feed-forward layers, time-frequency dense local convolutions, and gated units to capture fine-grained local details. Built upon a U-Net-like encoder-decoder structure with only 16 output channels in the encoder, LORT processes noisy inputs through multi-resolution T-MSA modules using alternating downsampling and upsampling operations. The enhanced magnitude and phase spectra are decoded independently and optimized through a composite loss function that jointly considers magnitude, complex, phase, discriminator, and consistency objectives. Experimental results on the VCTK+DEMAND and DNS Challenge datasets demonstrate that LORT achieves competitive or superior performance to state-of-the-art (SOTA) models with only 0.96M parameters, highlighting its effectiveness for real-world speech enhancement applications with limited computational resources.

语音增强轻量模型注意力机制卷积网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。