数据几何让时间可被推断,无需显式时间条件也能高效训练流匹配模型。
What Time Is It? How Data Geometry Makes Time Conditioning Optional for Flow Matching

- 利用高维数据的几何结构,从噪声观测中恢复时间信息。
- 时间盲训练的误差主要来自耦合方式,而非忽略时间本身。
- 实验证明调整数据配对比添加时间条件影响更大。
近期研究发现,流匹配模型可在无显式时间条件的情况下训练,挑战了传统观点——即插值时间用于区分速度目标。我们分解时间盲流匹配损失,识别出两个不可消除的误差源:耦合方差(由噪声与数据点配对引起的速度目标歧义)和时间盲差距(忽略时间带来的额外误差)。该差距表明时间盲训练严格更难,这加剧了为何时间盲模型在实践中表现良好的谜题。我们通过证明高维数据的几何结构使时间能直接从噪声观测中识别而解决这一矛盾:当数据集中在k维子空间时,可通过正交方向上的噪声插值统计结构恢复时间;在稀疏协方差模型下,单次观测即可以 $O(1/\\/sqrt{d-k})$ 的速率估计时间 $t$。因此,我们证明时间盲差距相对于耦合方差可忽略不计。我们在真实数据集上验证了可识别性,并表明改变耦合方式对损失和生成质量的影响远大于移除时间条件,在CIFAR-10、CelebA-HQ和FFHQ上均成立。结果解释了时间盲流匹配为何有效,指出实际关键调控因素是耦合选择而非显式时间条件。
原文摘要 · Abstract (English)
Recent work has shown that models flow matching models can be trained without explicit time conditioning, challenging the standard view that the interpolation time is needed to disambiguate velocity targets. But why should a time-blind model work at all? Decomposing the time-blind flow matching loss, we identify two sources of irreducible error: a coupling variance, which arises from ambiguous velocity targets induced by how noise and data points are paired, and the time-blindness gap, which is the additional error caused by ignoring time. This gap shows that time-blind training is strictly harder than conventional training, reinforcing the puzzle that time-blind models work so well in practice. We resolve this tension by showing that the geometry of high-dimensional data makes time identifiable directly from noisy observations. When data concentrates near a $k$-dimensional subspace, time can be recovered from the statistical structure of noisy interpolants in directions orthogonal to the data; under a spiked-covariance model, this yields a closed-form estimator that recovers $t$ from a single observation $z$ at rate $O(1/\sqrt{d-k})$ for ambient dimension $d$. As a consequence, we prove that the time-blindness gap is asymptotically negligible relative to the coupling variance. We empirically demonstrate our identifiability result on real-world data and show that changing the coupling has a much larger effect on loss and sample quality than removing time conditioning across CIFAR-10, CelebA-HQ, and FFHQ. These results explain why time-blind flow matching works and show that the main practical lever is the choice of coupling, not explicit time conditioning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。