arXiv:2606.24035eess.AS2026-06

揭示扩散模型语音增强在噪声失配时性能突变的根源

A Variational-Flow Analysis of Diffusion-Based Speech Enhancement under Noise-Power Mismatch

论文配图:A Variational-Flow Analysis of Diffusion-Based Speech Enhancement under Noise-Power Mismatch
图 1 · 摘自论文原文
  • 通过变分流分析定位性能突变源于确定性预测器
  • 发现噪声强度变化时输出不光滑,与得分网络无关
  • 为实际推理中的数值采样器提供理论框架

基于扩散的语音增强模型将确定性预测器与学习得到的得分网络结合,在训练时噪声幅度处表现出显著的非光滑退化(“拐点”)。本文提出路径型变分流分析,将该非光滑性定位至预测器阶段。核心结论是参数敏感性的精确分解:∂σ^(M)/∂M = K(M) · ∂C_M/∂M,其中K(M)为反向轨迹上得分雅可比的连续矩阵函数,C_M = Π(y^(M))为预测器输出。在反向过程流满足得分雅可比连续性、条件雅可比连续性及K非退化三个假设下,M↦σ^(M)在M*处非C^1当且仅当M↦Π(y^(M))在M*处非C^1。该分析被扩展至实际推理中使用的有限步欧拉-马鲁亚姆采样器。假设可转化为具体实验方案;本文给出该程序并呈现变分结构,实证验证留待配套实验报告。

原文摘要 · Abstract (English)

Diffusion-based speech enhancement architectures that pair a deterministic predictor with a learned score network, exhibit a sharp non-smooth transition (``kink'') in the SI-SDR degradation curve at the training-time noise amplitude. We give a pathwise variational-flow analysis that localizes this non-smoothness to the predictor stage. The central identity is an exact factorization of the parametric sensitivity, $\partial \sig^{(M)} / \partial M = K(M) \cdot \partial C_M / \partial M$, where $K(M)$ is a continuous matrix-valued functional of the score Jacobian along the reverse trajectory and $C_M = Π(y^{(M)})$ is the predictor output. Under three hypotheses on the reverse-process flow (score-Jacobian continuity, conditioning-Jacobian continuity, non-degeneracy of $K$), failure of $M \mapsto \sig^{(M)}$ to be $C^1$ at $M^\ast$ holds if and only if $M \mapsto Π(y^{(M)})$ fails to be $C^1$ at $M^\ast$. We extend the localization to the finite-step Euler--Maruyama sampler actually run at inference. The hypotheses translate into a concrete experimental program; this paper specifies the program and presents the variational structure. The empirical validation is deferred to a companion experimental report.

语音增强扩散模型非光滑性变分分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。