arXiv:2606.20001eess.AS2026-06

无需时间步信息,直接通过空间关系实现语音增强

Time-Unconditional Generative Speech Enhancement via Autonomous Rectified Flow

  • 用线性插值证明目标向量场天然与时间无关
  • 摒弃显式时间编码,仅凭当前状态与噪声观测推断去噪方向
  • 提升生成质量、鲁棒性及推理效率,适合追求高效语音增强的场景

大多数生成式语音增强方法依赖显式的时间步嵌入进行时序条件建模。本文提出自主修正流(Autonomous Rectified Flow)框架,挑战了这种时序条件的必要性。通过线性插值路径,我们证明目标向量场本质上是时不变的。进一步设计了无时间条件网络,消除显式时间步信息,仅依据当前状态与噪声观测之间的空间关系推断去噪方向。预测该目标向量场等价于建模噪声分布。通过避免对时序轨迹的过拟合,所提出的自主设计显著提升了生成质量、鲁棒性及推理效率。

原文摘要 · Abstract (English)

Most generative speech enhancement methods rely on explicit time-step embeddings for temporal conditioning. In this paper, we propose the Autonomous Rectified Flow framework, which challenges the necessity of such conditioning. Using a linear interpolation path, we show that the target vector field is inherently time-invariant. We further introduce a time-unconditional network that eliminates explicit time-step information and infers the denoising direction solely from the spatial relationship between the current state and the noisy observation. Predicting this target vector field is equivalent to modeling the noise distribution. By avoiding overfitting to temporal trajectories, the proposed autonomous design significantly improves generation quality, robustness, and inference efficiency.

语音增强生成模型扩散模型时序无关

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。