arXiv:2412.14435cs.LGcs.AI2024-12被引 14

研究发现选对数据集能让模型表现‘看起来’更好,提醒评估要全面。

Cherry-Picking in Time Series Forecasting: How to Select Datasets to Make Your Model Shine

  • 通过分析多个基准数据集,发现选特定数据集会夸大模型性能。
  • 仅选4个数据集时,46%的模型能被宣称最优,77%进前三。
  • 深度学习模型更易受数据选择影响,经典方法更稳定。

时间序列预测的重要性推动了新方法的持续研究,但多数方法依赖实证实验声称精度优越。然而,实验设置的局限性引发对其结果可靠性与泛化能力的担忧。本文聚焦于数据集选择偏差这一关键问题,特别是“樱桃采摘”现象对方法性能评估的影响。通过对多样化的基准数据集进行实证分析,结果表明:选择性使用数据集会显著扭曲方法的性能表现,过度夸大其有效性。具体而言,仅选取4个数据集(多数研究报告的数量)时,46%的方法可被标为最佳,77%可进入前三名。此外,近期基于深度学习的方法对数据集选择高度敏感,而传统方法表现出更强鲁棒性。最后,当在基准子集上验证算法时,将测试数据集数量从3个增加到6个,可使错误识别最佳算法的风险降低约40%。研究强调需建立更全面的评估框架,以更真实反映实际场景,保障预测方法的稳健性和可靠性。

原文摘要 · Abstract (English)

The importance of time series forecasting drives continuous research and the development of new approaches to tackle this problem. Typically, these methods are introduced through empirical studies that frequently claim superior accuracy for the proposed approaches. Nevertheless, concerns are rising about the reliability and generalizability of these results due to limitations in experimental setups. This paper addresses a critical limitation: the number and representativeness of the datasets used. We investigate the impact of dataset selection bias, particularly the practice of cherry-picking datasets, on the performance evaluation of forecasting methods. Through empirical analysis with a diverse set of benchmark datasets, our findings reveal that cherry-picking datasets can significantly distort the perceived performance of methods, often exaggerating their effectiveness. Furthermore, our results demonstrate that by selectively choosing just four datasets - what most studies report - 46% of methods could be deemed best in class, and 77% could rank within the top three. Additionally, recent deep learning-based approaches show high sensitivity to dataset selection, whereas classical methods exhibit greater robustness. Finally, our results indicate that, when empirically validating forecasting algorithms on a subset of the benchmarks, increasing the number of datasets tested from 3 to 6 reduces the risk of incorrectly identifying an algorithm as the best one by approximately 40%. Our study highlights the critical need for comprehensive evaluation frameworks that more accurately reflect real-world scenarios. Adopting such frameworks will ensure the development of robust and reliable forecasting methods.

时间序列评估偏差数据集选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。