提出新生成框架,提升语音增强的训练效率与音质。
Target matching based generative model for speech enhancement
- 将语音增强转为目标信号估计,去除随机成分提升稳定性。
- 采用逻辑均值与桥接方差调度,信噪比轨迹更优。
- 设计新音频扩散主干,显著降低复杂度,加速推理。
生成模型中扰动信号的均值和方差调度设计是核心挑战。尽管基于分数的和基于薛定谔桥的模型需精心选择随机微分方程以推导对应调度,流模型则通过向量场匹配解决此问题,但该策略常引入幻觉伪影,并因向量场中潜在的随机成分导致训练与推理效率低下。此外,广泛使用的扩散主干NCSN++计算复杂度较高。为此,我们提出一种新型基于目标的生成框架,既提升了均值/方差调度设计的灵活性,又优化了训练与推理效率。具体地,我们将生成式语音增强任务重构为目标信号估计问题,消除训练损失中的随机成分,从而实现更稳定高效的训练与推理。同时,采用逻辑均值调度与桥接方差调度,相比多种常用调度方案,可获得更优的信噪比演化轨迹,实现更高效的扰动策略。此外,我们提出一种新型音频扩散主干,通过显式建模长期帧相关性和跨频带依赖关系,显著优于NCSN++的效率表现。
原文摘要 · Abstract (English)
The design of mean and variance schedules for the perturbed signal is a fundamental challenge in generative models. While score-based and Schrödinger bridge-based models require careful selection of the stochastic differential equation to derive the corresponding schedules, flow-based models address this issue via vector field matching. However, this strategy often leads to hallucination artifacts and inefficient training and inference processes due to the potential inclusion of stochastic components in the vector field. Additionally, the widely adopted diffusion backbone, NCSN++, suffers from high computational complexity. To overcome these limitations, we propose a novel target-based generative framework that enhances both the flexibility of mean/variance schedule design and the efficiency of training and inference processes. Specifically, we eliminate the stochastic components in the training loss by reformulating the generative speech enhancement task as a target signal estimation problem, which therefore leads to more stable and efficient training and inference processes. In addition, we employ a logistic mean schedule and a bridge variance schedule, which yield a more favorable signal-to-noise ratio trajectory compared to several widely used schedules and thus leads to a more efficient perturbation strategy. Furthermore, we propose a new diffusion backbone for audio, which significantly improves the efficiency over NCSN++ by explicitly modeling long-term frame correlations and cross-band dependencies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。