arXiv:2607.19383stat.APcs.LG2026-07

趋势强弱决定生成模型能否胜出,可提前判断是否用它预测。

Trend strength predicts when generative foundation models win: a power-controlled benchmark, a mechanism, and an actionable selection rule

  • 用趋势强度预判生成模型是否优于传统方法
  • 低趋势序列上生成模型胜率78%,高趋势仅44%
  • 无需试错,凭训练数据趋势即可选模型

预训练生成式基础模型将预测视为从学习到的预测分布中进行条件生成,并实现零样本预测未见序列。我们建立三项成果,使报告的成功转化为可操作、机制化的理解。第一(正向基准结果):在涵盖36个序列、19个数据集、覆盖全范围的STL趋势强度(F_T ∈ [0.17, 1.00])的1728次滚动起点预测中,零样本Chronos模型显著优于四种强基线——漂移、季节性朴素、Theta和加法Holt-Winters/ETS——最佳平均MASE为1.187(Theta: 1.337,ETS: 1.656;Friedman卡方=46.08,p=8.75e-09;Holm校正后的Wilcoxon p ≤ 0.015,对所有基线均显著;Nemenyi临界差异将其与经典模型分离)。第二(新颖量化机制):在已知趋势生成过程的受控合成实验中揭示原因——胜利并非来自更好趋势外推。当真实趋势为线性、阻尼或指数时,加法ETS的斜率追踪比分别为1.02、1.34、0.98,而Chronos系统性低估,表现为趋势收缩估计器(比例0.80、0.49、0.36)。第三(可操作选择规则):因优势源于收缩效应,其可由趋势强度单独预测——在低趋势序列中获胜率达78%,高趋势仅为44%;对低趋势子组,其优势显著(0.982 vs. 1.671,p=0.002),高趋势则持平(p=0.18)。趋势强度可仅从训练上下文计算,是部署基础模型前的实用先验指标。此外还发现校准不足(80%区间覆盖率为0.77)。

原文摘要 · Abstract (English)

Pretrained generative foundation models cast forecasting as conditional generation from a learned predictive distribution and forecast unseen series zero-shot. We establish three results that turn their reported success into an actionable, mechanistic understanding. First (a positive benchmark result): on a power-controlled study of 1728 rolling-origin forecasts over 36 series from 19 datasets spanning the full range of STL trend strength (F_T in [0.17, 1.00]), a zero-shot Chronos model significantly outperforms four strong classical baselines -- drift, seasonal-naive, Theta, and additive Holt-Winters/ETS -- with the best mean MASE (1.187 vs. Theta 1.337, ETS 1.656; Friedman chi^2 = 46.08, p = 8.75e-09; Holm-corrected Wilcoxon p <= 0.015 against every baseline; a Nemenyi critical difference separating it from the classical pack). Second (a novel, quantified mechanism): a controlled synthetic experiment with a known trend-generating process shows why -- and reveals that the win does not come from better trend extrapolation. When the true trend is linear, damped, or exponential, additive ETS tracks the slope (slope-tracking ratio 1.02, 1.34, 0.98) whereas Chronos systematically under-extrapolates, behaving as a trend-shrinkage estimator (ratio 0.80, 0.49, 0.36). Third (an actionable selection rule): because the advantage is a shrinkage effect, it is predictable from trend strength alone -- the generative model wins 78% of low-trend series but only 44% of high-trend ones, and its edge over ETS is significant on the low-trend stratum (0.982 vs. 1.671, p = 0.002) yet a tie on the high-trend stratum (p = 0.18). Trend strength, computable before forecasting from the training context alone, is therefore a practical a-priori indicator of when to deploy a foundation model. We additionally document a calibration shortfall (80% intervals cover 0.77).

时间序列生成模型趋势预测模型选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。