arXiv:2609.04525eess.AScs.SD2026-09

用判别式表征替代时间条件,让生成修复更精准

Discriminative Flow Matching: Beyond Time-Conditioning in Generative Restoration via Flow-State Representations

论文配图:Discriminative Flow Matching: Beyond Time-Conditioning in Generative Restoration via Flow-State Representations
图 1 · 摘自论文原文
  • 用判别模型学习的表征作为生成状态,代替传统时间变量
  • 在语音增强和图像去噪任务中性能超越基线方法
  • 适合需要自适应推理的生成修复场景

现有条件流匹配(CFM)通过显式插值坐标(常被理解为时间)描述生成过程,但恢复任务中不同样本的退化程度差异大,相同插值位置的样本可能在退化水平、与目标分布距离和修复难度上显著不同。本文研究判别训练模型所学表征是否能有效描述生成状态。通过系统的潜在空间分析发现,这些表征按退化严重程度组织,并沿一致轨迹向干净数据流形演化。基于此提出判别流状态假设:判别表征编码了生成修复的传输状态。据此提出判别流匹配,将流匹配速度场条件于判别流状态表示而非显式时间坐标。在语音增强和图像去噪任务上的实验表明,该表征能刻画修复进度,支持自适应推理,并持续优于CFM及扩散类基线方法。结果表明,判别表征是显式时间条件的有效状态感知替代方案,为判别模型与流匹配生成建模的关系提供了新视角。

原文摘要 · Abstract (English)

Existing Conditional Flow Matching (CFM) formulations describe transport progress using an explicit interpolation coordinate, commonly interpreted as time, assuming that a single global variable adequately represents a sample's position along the generative trajectory. In restoration tasks, however, transport progress is sample-dependent because the initial distribution may exhibit varying statistical dependencies with the target distribution. Thus, samples at the same interpolation coordinate can differ substantially in degradation level, distance to the target distribution, and restoration difficulty. We investigate whether signal representations learned by discriminatively trained models provide a meaningful description of generative transport state in CFM-based restoration. Through systematic latent-space analysis, we show that discriminative representations organize according to degradation severity and follow a consistent trajectory toward the clean-data manifold during generation. Motivated by these observations, we introduce the Discriminative Flow-State Hypothesis, which posits that discriminative representations encode a transport state governing generative restoration. Based on this hypothesis, we propose Discriminative Flow Matching, which conditions the Flow-Matching velocity field on Discriminative Flow-State Representations rather than explicit time coordinates. Experiments on speech enhancement and image denoising show that these representations characterize restoration progress, enable adaptive inference, and consistently outperform CFM and diffusion-related baselines. Our findings suggest that discriminative representations provide an effective state-aware alternative to explicit time conditioning and offer a novel perspective on the relationship between discriminative and CFM-based generative modeling.

生成修复流匹配判别表征自适应推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。