arXiv:2608.21188eess.AS2026-08

通过动态调整网络宽度,让语音增强扩散模型更高效

SlimDiffuSE: Towards Efficient Diffusion-Based Speech Enhancement using Slimmable Networks

  • 用可裁剪网络在生成过程中自适应调整宽度
  • 计算量降低87.5%仍保持高语音质量
  • 适合资源受限场景下的实时语音增强

扩散模型在语音增强领域崭露头角,已在多个基准数据集上达到顶尖性能。但其主要缺陷是生成数据需多次调用大型神经网络,导致整体计算复杂度高。本文提出一种可裁剪的扩散模型,在生成过程中自适应调整网络宽度以降低计算开销。通过贪心搜索算法优化网络宽度调度策略,该方法在显著降低计算复杂度的同时,性能与基线扩散模型相当。值得注意的是,该方法在不明显降低客观指标(如语音质量感知评价PESQ和信号失真比SI-SDR)的前提下,计算复杂度最高可降低87.5%。

原文摘要 · Abstract (English)

Diffusion-based models are emerging in the speech enhancement domain and are achieving state-of-the-art performance across various benchmark datasets. A major downside of diffusion models is that data generation requires many evaluations of a typically large neural network, which results in high overall complexity. In this work, we propose a slimmable diffusion model that employs adaptive network widths throughout the data generation process to reduce computational cost. By using a greedy search algorithm to optimize the network width schedule, our method achieves performance comparable to baseline diffusion models with significantly reduced computational complexity. Notably, our approach reduces the computational complexity by up to $87.5\%$ without a significant drop in objective metrics, such as perceptual evaluation of speech quality (PESQ) and SI-SDR.

扩散模型语音增强高效计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。