提出新评估框架,让时间序列预测模型对比更公平。
Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting
- 基于谱相干性设计评分体系,量化数据可预测难度。
- 发现任务可预测性随时间剧烈变化,且复杂模型未必更优。
- 适合关注模型真实性能与可解释性的研究者使用。
在时间序列预测模型日益复杂的背景下,进展常以基准排行榜上的微小提升来衡量。然而,这一方法存在根本缺陷:标准评估指标将模型表现与数据固有的不可预测性混为一谈。为此,我们提出一种基于谱相干性的可预测性对齐诊断框架。该框架包含两个核心贡献:一是计算高效($O(N\log N)$)的谱相干可预测性(SCP)得分,用于量化特定预测实例的内在难度;二是频率分辨的线性利用比(LUR),精确测量模型对数据中线性可预测信息的利用效率。通过验证,我们首次系统揭示了‘可预测性漂移’现象,即任务难度随时间显著变化;同时发现关键架构权衡:复杂模型在低可预测性数据上更优,而线性模型在高可预测任务中表现极佳。我们倡导从单一平均分转向更深入、可预测性感知的评估范式,以实现更公平的模型比较和对模型行为的深层理解。
原文摘要 · Abstract (English)
In the era of increasingly complex AI models for time series forecasting, progress is often measured by marginal improvements on benchmark leaderboards. However, this approach suffers from a fundamental flaw: standard evaluation metrics conflate a model's performance with the data's intrinsic unpredictability. To address this pressing challenge, we introduce a novel, predictability-aligned diagnostic framework grounded in spectral coherence. Our framework makes two primary contributions: the Spectral Coherence Predictability (SCP), a computationally efficient ($O(N\log N)$) and task-aligned score that quantifies the inherent difficulty of a given forecasting instance, and the Linear Utilization Ratio (LUR), a frequency-resolved diagnostic tool that precisely measures how effectively a model exploits the linearly predictable information within the data. We validate our framework's effectiveness and leverage it to reveal two core insights. First, we provide the first systematic evidence of "predictability drift", demonstrating that a task's forecasting difficulty varies sharply over time. Second, our evaluation reveals a key architectural trade-off: complex models are superior for low-predictability data, whereas linear models are highly effective on more predictable tasks. We advocate for a paradigm shift, moving beyond simplistic aggregate scores toward a more insightful, predictability-aware evaluation that fosters fairer model comparisons and a deeper understanding of model behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。