arXiv:2510.13654cs.LGcs.AI2025-10被引 9

TSFM评估存在信息泄露,可能导致结果虚高。

Rethinking Evaluation in the Era of Time Series Foundation Models: (Un)known Information Leakage Challenges

  • 发现两类信息泄露:数据集复用导致样本重叠,时间相关序列重叠。
  • 忽略泄露会使预测性能评估严重偏高,无法反映真实效果。
  • 呼吁建立更严谨的评估方法,保障模型评测可信度。

时间序列基础模型(TSFMs)为时间序列预测带来新范式,可实现无需任务特定训练的零样本预测。然而,与大语言模型(LLMs)类似,其评估面临挑战:随着训练语料规模扩大,难以确保基准测试集的完整性。对现有TSFM评估研究的分析揭示了两类信息泄露:(1) 因数据集多用途复用导致的训练-测试样本重叠;(2) 相关训练与测试序列间的时间重叠。忽视这些泄露会导致性能评估过于乐观,无法泛化至真实场景。因此,我们主张开发新型评估方法,规避已有在LLM及经典时间序列基准中观察到的陷阱,并呼吁研究界采用严谨方法以维护TSFM评估的完整性。

原文摘要 · Abstract (English)

Time Series Foundation Models (TSFMs) represent a new paradigm for time-series forecasting, promising zero-shot predictions without the need for task-specific training or fine-tuning. However, similar to Large Language Models (LLMs), the evaluation of TSFMs is challenging: as training corpora grow increasingly large, it becomes difficult to ensure the integrity of the test sets used for benchmarking. An investigation of existing TSFM evaluation studies identifies two kinds of information leakage: (1) train-test sample overlaps arising from the multi-purpose reuse of datasets and (2) temporal overlap of correlated train and test series. Ignoring these forms of information leakage when benchmarking TSFMs risks producing overly optimistic performance estimates that fail to generalize to real-world settings. We therefore argue for the development of novel evaluation methodologies that avoid pitfalls already observed in both LLM and classical time-series benchmarking, and we call on the research community to adopt principled approaches to safeguard the integrity of TSFM evaluation.

时间序列模型评估信息泄露

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。