arXiv:2510.11395eess.AS2025-10中稿 · ICASSP2026被引 2

根据语音质量动态调整模型计算量,提升轻量化语音增强效率

Dynamically Slimmable Speech Enhancement Network with Metric-Guided Training

  • 用门控机制让模型按输入质量动态启用不同组件
  • 在保持性能的前提下,平均减少27%计算量
  • 适合资源受限设备的实时语音增强场景

为降低轻量级语音增强模型的复杂度,本文提出基于门控的动态瘦身网络(DSN)。DSN包含静态与动态两部分,针对常用的分组循环单元、多头注意力、卷积和全连接层设计独立动态结构。策略模块根据输入信号质量在帧级自适应控制动态组件的使用,调节计算负载。进一步提出度量引导训练(MGT),显式指导策略模块评估语音质量。实验表明,DSN在仪器度量上达到当前最优轻量基线水平,平均仅需其73%的计算量。动态组件使用率评估显示,MGT-DSN能根据输入信号失真程度合理分配资源。

原文摘要 · Abstract (English)

To further reduce the complexity of lightweight speech enhancement models, we introduce a gating-based Dynamically Slimmable Network (DSN). The DSN comprises static and dynamic components. For architecture-independent applicability, we introduce distinct dynamic structures targeting the commonly used components, namely, grouped recurrent neural network units, multi-head attention, convolutional, and fully connected layers. A policy module adaptively governs the use of dynamic parts at a frame-wise resolution according to the input signal quality, controlling computational load. We further propose Metric-Guided Training (MGT) to explicitly guide the policy module in assessing input speech quality. Experimental results demonstrate that the DSN achieves comparable enhancement performance in instrumental metrics to the state-of-the-art lightweight baseline, while using only 73% of its computational load on average. Evaluations of dynamic component usage ratios indicate that the MGT-DSN can appropriately allocate network resources according to the severity of input signal distortion.

语音增强动态网络轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。