arXiv:2507.04879eess.AScs.LG2025-07中稿 · publication at the…被引 3

让语音增强模型按需瘦身,省算力还保效果。

Adaptive Slimming for Scalable and Efficient Speech Enhancement

  • 动态调整模型复杂度,输入不同则用不同规模
  • 平均只用10%算力,效果却超静态25%模型
  • 适合移动端等资源受限场景的语音增强应用

语音增强(SE)在语音识别、实时通信和助听器等场景中至关重要。然而,在资源受限设备上部署时,性能与效率常需静态权衡。本文为主流的DEMUCS架构引入动态瘦身机制,使其可自适应调节利用率(UF),在不增加存储成本的前提下,实现多尺寸模型的灵活切换。通过端到端训练的路由子网络,系统根据输入自动选择最优UF,避免冗余计算。实验表明,该方案在帕累托意义上优于单一固定UF模型。当模型平均仅使用10%容量时,其语音质量达到甚至超过静态25%利用率的基准,同时减少29%的乘加操作(MACs)。

原文摘要 · Abstract (English)

Speech enhancement (SE) enables robust speech recognition, real-time communication, hearing aids, and other applications where speech quality is crucial. However, deploying such systems on resource-constrained devices involves choosing a static trade-off between performance and computational efficiency. In this paper, we introduce dynamic slimming to DEMUCS, a popular SE architecture, making it scalable and input-adaptive. Slimming lets the model operate at different utilization factors (UF), each corresponding to a different performance/efficiency trade-off, effectively mimicking multiple model sizes without the extra storage costs. In addition, a router subnet, trained end-to-end with the backbone, determines the optimal UF for the current input. Thus, the system saves resources by adaptively selecting smaller UFs when additional complexity is unnecessary. We show that our solution is Pareto-optimal against individual UFs, confirming the benefits of dynamic routing. When training the proposed dynamically-slimmable model to use 10% of its capacity on average, we obtain the same or better speech quality as the equivalent static 25% utilization while reducing MACs by 29%.

语音增强动态模型高效推理轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。