arXiv:2608.09692cs.LGcs.AI2026-08

发现时间序列生成模型评估中零值分布偏差会严重误导结论

Evaluating Generative Time-Series Models on Data with Point Masses

  • 提出新评估协议,避免因窗口零值比例与数据集差异导致误判
  • 自回归漏斗模型在6个数据集上表现优于条件流模型最高153倍
  • 不同评估指标排序不一致,揭示模型选择高度依赖评价方式

许多用于评估生成式时间序列模型的数据集中存在大量零值(如无降雨、无订单),这使得标准滚动原点评估协议可能在零值占比与整体数据集显著不同的窗口上进行测试:一个数据集零值占比42%,但评估窗口仅为13%;另一个为47%对5%。这一问题非表面现象,甚至逆转了我们原有的结论,使最强模型看似反面教材。我们提出一种控制实验,使CRPS指标在构造上保持不变,但破坏时间耦合性,从而精确衡量时间依赖对统计量的贡献。在五个种子下对七种模型进行匹配协议评测发现,自回归漏斗模型在六个数据集中的五个上胜过条件流模型,性能差距高达153倍;而流模型自身发生率统计在不同训练种子间波动达62%,所有基线模型均为确定性。此外,模型排序随五种不同发生率统计量变化,其中两种未共享构造的指标之间一致性最低。

原文摘要 · Abstract (English)

Many of the series that generative time-series models are benchmarked on place a large probability mass on a single value --- it does not rain, no ride is requested, no part is ordered. We report what happens when such data is evaluated carefully. First, the standard rolling-origin protocol can score a model on a window whose atom structure bears no resemblance to the dataset: on one benchmark the dataset is $42\%$ zeros and the evaluation windows are $13\%$, on another $47\%$ against $5\%$. This is not a cosmetic problem --- it reversed one of our own conclusions, turning the strongest occurrence model in our study into what looked like a cautionary tale. Second, we give a control in which CRPS is invariant \emph{by construction} while the temporal coupling is destroyed, which measures exactly how much that coupling contributes to a chosen statistic. Third, benchmarking seven models on a matched protocol over five seeds, an autoregressive hurdle beats a conditional flow on five of six datasets, by up to a factor of $153$, while the flow's own occurrence statistics vary by up to $62\%$ across training seeds and every baseline is deterministic. Finally, the model ordering is not the same under five different occurrence statistics, and the two that do not share a construction agree with each other least.

时间序列生成评估偏差零值数据模型比较

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。