arXiv:2608.20980cs.LGstat.ML2026-08

发现主流时空预测数据集存在结构偏差,线性模型竟胜过图神经网络。

A Critical Audit of Spatiotemporal Forecasting Benchmark Datasets and Baselines

论文配图:A Critical Audit of Spatiotemporal Forecasting Benchmark Datasets and Baselines
图 1 · 摘自论文原文
  • 用经典时间序列方法分析基准数据集,揭示其内在统计特性。
  • 发现一阶差分导致数据结构偏倚,使线性模型表现异常出色。
  • 建议重构评估流程,适合做模型公平对比的研究者参考。

图神经网络(GNNs)常用于具有空间图结构的多变量时间序列短时预测。尽管存在多种替代数据集,该领域的方法创新仍主要基于少数几个基准数据集,如Chickenpox、PedalMe、WikiMaths、METR-LA和PEMS-BAY。评估协议包含从历史均值到传统机器学习方法的基线模型,这些基线常表现出与GNN相当甚至更优的性能。本文通过经典时间序列方法重新审视这些基准数据集,揭示为何无空间感知的线性模型能成为强竞争者,进一步质疑了这些广泛使用数据集的判别可靠性。我们的统计分析提供了一套识别显著时空相关性的工具,并发现一阶差分数据集引入了结构性偏差。因此我们建议减少对这些数据集的过度依赖,转而倡导更严格的统计评估。通过将分析结果应用于一个简单混合模型,展示了该方法如何推动GNN模型的新设计思路。

原文摘要 · Abstract (English)

Graph neural networks (GNNs) are routinely employed for short-range forecasting on multivariate time series with a spatial graph structure. Despite the availability of many alternative datasets, method innovations within this domain are predominantly assessed against a rather limited set of benchmark datasets, most notably Chickenpox, PedalMe, WikiMaths, METR-LA, and PEMS-BAY. The evaluation protocols contain baselines spanning from historical averages to classical machine learning approaches. These baselines often show competitive performance compared to GNNs. In the present work, we take a step back and analyse the benchmark datasets via classical time series methods to uncover why spatially-unaware linear models pose a stronger competitor than previously reported, casting further doubt on the discriminative reliability of the aforementioned widely adopted datasets. Our statistical analysis provides a toolset for identifying significant spatial and temporal correlations, while revealing a structural bias introduced by first-order differenced datasets. We therefore recommend reducing the over-reliance on such datasets for method comparison, and instead advocate for more rigorous statistical evaluation. By applying the results of our analysis to a simple hybrid model, we show how our methodology can lead to novel ways of developing GNN models

时空预测基准测试图神经网络数据偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。