提出多维评估框架,打破时间序列预测仅看误差的旧游戏。
Are We Winning the Wrong Game? Revisiting Evaluation Practices for Long-Term Time Series Forecasting
- 用统计保真、结构一致性和决策相关性重构评估维度
- 发现现有误差指标无法反映真实场景中的趋势与季节稳定性
- 适合关注实际应用价值的研究者和工业界从业者
长期时间序列预测(LTSF)被视为数据挖掘与机器学习的核心挑战。当前研究逐渐演变为以基准测试为驱动的“游戏”,模型优劣主要依据均方误差(MSE)和平均绝对误差(MAE)等点对点误差指标的微小降低来评判。在少数经典数据集和固定预测时长下,进展通过排行榜式表格呈现,数值越低即代表成功。在这种以度量为中心的体系中,被测量的成为被优化的,误差的边际下降成为进步的主要货币。我们指出,这种度量主导的模式不仅不完整,且在结构上与预测的总体目标错位。在真实场景中,预测更重视保持时间结构、趋势稳定性、季节一致性、对制度突变的鲁棒性以及支持下游决策过程。单纯优化聚合点对点误差,并不必然意味着建模了这些结构性特征。因此,排行榜提升可能更多反映对基准配置的适应,而非对时间动态的深层理解。本文重新审视LTSF评估这一基础性问题:何谓预测进展的衡量?我们提出融合统计保真、结构一致性和决策相关性的多维评估视角。通过挑战当前的度量单一化,旨在引导关注从赢得排行榜转向实现有意义、上下文感知的预测能力。
原文摘要 · Abstract (English)
Long-term time series forecasting (LTSF) is widely recognized as a central challenge in data mining and machine learning. LTSF has increasingly evolved into a benchmark-driven ''GAME,'' where models are ranked, compared, and declared state-of-the-art based primarily on marginal reductions in aggregated pointwise error metrics such as MSE and MAE. Across a small set of canonical datasets and fixed forecasting horizons, progress is communicated through leaderboard-style tables in which lower numerical scores define success. In this GAME, what is measured becomes what is optimized, and incremental error reduction becomes the dominant currency of advancement. We argue that this metric-centric regime is not merely incomplete, but structurally misaligned with the broader objectives of forecasting. In real-world settings, forecasting often prioritizes preserving temporal structure, trend stability, seasonal coherence, robustness to regime shifts, and supporting downstream decision processes. Optimizing aggregate pointwise error does not necessarily imply modeling these structural properties. As a result, leaderboard improvement may increasingly reflect specialization in benchmark configurations rather than a deeper understanding of temporal dynamics. This paper revisits LTSF evaluation as a foundational question in data science: what does it mean to measure forecasting progress? We propose a multi-dimensional evaluation perspective that integrates statistical fidelity, structural coherence, and decision-level relevance. By challenging the current metric monoculture, we aim to redirect attention from winning benchmark tables toward advancing meaningful, context-aware forecasting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。