arXiv:2607.24662cs.LGcs.SI2026-07

时间图生成模型在部署时会因分布漂移而失效,且无法仅靠观测数据修复。

When Can You Correct Distribution Drift in Temporal Graph Generation? A Sharpening--Drift Tension and an Impossibility for Observation-Based Correction

  • 通过分解掩码流匹配损失,揭示漂移不可逆的熵增本质
  • 实验显示误差下限随采样预算变化极小,漂移使误差提升2.2至34.3倍
  • 基于观测的修正方法效果极差,趋势外推反而不如不修正

时间图生成模型在一段演化网络上训练后部署到下一阶段,性能会严重下降。我们证明这种退化是可推导、普遍存在的,且无法仅从观测中修复。掩码流匹配损失可精确分解,无需独立性假设,包含不可约熵项和一个导数为正的发散项,该导数对应于训练期罕见但部署期常见的结构,其训练概率趋近零时发散加剧。实证发现该权衡呈幂律关系,指数为-0.605(R²=0.9977)。漂移提升了采样器的误差下限,但不影响达到该下限所需的采样步数:在七个高统计功效条件下,漂移周期的边际误差在采样预算扩大50倍时仅变化最多6%,而下限比同期下限高出2.2至34.3倍。尽管部署周期可观测,修正看似可行,实则不然。我们证明任何关于历史观测可测的修正器,至少保留其所追踪统计量的条件方差;趋势外推优于依赖最近观测的条件是μ² > v(1−2ρ),但两者前提均不成立——漂移无趋势且均值回归,一步创新量与漂移量相当。预言机可消除60%误差,最优观测修正仅恢复其中5.7%,而趋势外推严格劣于简单不修正。

原文摘要 · Abstract (English)

Generative models of temporal graphs are trained on one stretch of an evolving network and deployed on the next, and they degrade badly in the gap. We show this degradation is derivable, general, and not fixable from observations. The masked flow-matching loss decomposes exactly, with no independence assumption, into an irreducible entropy plus a divergence whose derivative along the training path is positive precisely for structures rare during training and common at deployment, diverging as their training probability goes to zero. Empirically the trade-off is a power law with exponent $-0.605$ ($R^2=0.9977$), and drift raises the sampler's error floor without changing how many steps reach it: across seven well-powered conditions the drift-period marginal error varies by at most $6\%$ over a $50\times$ range of sampling budgets, while the floor sits $2.2\times$ to $34.3\times$ above the in-period floor. Because the deployment period is observed, correction looks like a matter of measurement. It is not. We prove that any corrector measurable with respect to past observations leaves at least the conditional variance of the statistic it tracks, and that trend extrapolation beats trusting the last observation only when $μ^2>v(1-2ρ)$. Both premises are measurable and both go the wrong way: the drift is trendless and mean-reverting, with a one-step innovation as large as the drift itself. An oracle removes $60\%$ of the error, the best observation-based corrector recovers $5.7\%$ of that, and extrapolation is strictly worse than doing nothing clever.

时间图生成分布漂移生成模型不可修复性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。