提出无需EMA的语音增强扩散模型,提升清晰度与稳定性。
Do We Need EMA for Diffusion-Based Speech Enhancement? Toward a Magnitude-Preserving Network Architecture
- 采用幅度保持架构与时间依赖预处理,稳定训练过程。
- 不同跳跃连接设计使模型可预测噪声或纯净语音,性能互补。
- 实验证明短或无EMA反而提升语音质量,颠覆图像生成经验。
我们基于薛定谔桥框架研究扩散模型在语音增强中的应用,并将EDM2框架扩展至该场景。通过引入时间依赖的输入输出预处理以稳定训练,探索两种跳跃连接配置,使网络能够预测环境噪声或纯净语音。为控制激活值和权重幅度,采用幅度保持架构,并学习每个网络块中噪声输入的贡献以改善条件建模。进一步通过后训练近似不同EMA配置,分析指数移动平均(EMA)参数平滑的影响,发现与图像生成不同,短或缺失EMA始终带来更优的语音增强表现。在VoiceBank-DEMAND和EARS-WHAM数据集上的实验表明,该方法在信号失真比和感知评分上均具竞争力,两种跳跃连接变体表现出互补优势。这些发现为扩散模型在语音增强中的EMA行为、幅度保持及跳跃连接设计提供了新见解。
原文摘要 · Abstract (English)
We study diffusion-based speech enhancement using a Schrodinger bridge formulation and extend the EDM2 framework to this setting. We employ time-dependent preconditioning of network inputs and outputs to stabilize training and explore two skip-connection configurations that allow the network to predict either environmental noise or clean speech. To control activation and weight magnitudes, we adopt a magnitude-preserving architecture and learn the contribution of the noisy input within each network block for improved conditioning. We further analyze the impact of exponential moving average (EMA) parameter smoothing by approximating different EMA profiles post training, finding that, unlike in image generation, short or absent EMA consistently yields better speech enhancement performance. Experiments on VoiceBank-DEMAND and EARS-WHAM demonstrate competitive signal-to-distortion ratios and perceptual scores, with the two skip-connection variants exhibiting complementary strengths. These findings provide new insights into EMA behavior, magnitude preservation, and skip-connection design for diffusion-based speech enhancement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。