arXiv:2510.20273cs.LG2025-10NeurIPS被引 4

用可编程合成数据评估时序模型的真实能力,发现现有模型并非处处最优。

SynTSBench: Rethinking Temporal Pattern Learning in Deep Learning Models for Time Series

  • 通过可控合成数据拆解时序模式,系统测试模型对不同特征的学习能力。
  • 揭示模型在噪声和异常下的容忍阈值与恢复能力,量化其鲁棒性短板。
  • 对比模型预测与数学最优解,明确各模型性能边界,适合选型参考。

深度学习在时序预测中取得显著进展,但许多先进模型在真实场景中仍表现不佳,即使在标准基准数据集上表现优异。这一差距源于深度模型的黑箱特性及现有评估框架的局限性,难以提供清晰、量化的模型优劣分析。为此,我们提出合成数据驱动的评估范式 SynTSBench,通过可编程特征配置系统评估时序模型的核心建模能力。该框架隔离混杂因素,建立三个核心分析维度:(1) 时序特征分解与能力映射,系统评估模型对特定模式类型的学习能力;(2) 数据不规则下的鲁棒性分析,量化噪声容忍阈值与异常恢复能力;(3) 理论最优基准对比,为每种模式类型建立性能上限,实现模型预测与数学最优解的直接比较。实验表明,当前深度学习模型并未在所有类型的时序特征上接近最优基线。代码已公开于 https://github.com/TanQitai/SynTSBench。

原文摘要 · Abstract (English)

Recent advances in deep learning have driven rapid progress in time series forecasting, yet many state-of-the-art models continue to struggle with robust performance in real-world applications, even when they achieve strong results on standard benchmark datasets. This persistent gap can be attributed to the black-box nature of deep learning architectures and the inherent limitations of current evaluation frameworks, which frequently lack the capacity to provide clear, quantitative insights into the specific strengths and weaknesses of different models, thereby complicating the selection of appropriate models for particular forecasting scenarios. To address these issues, we propose a synthetic data-driven evaluation paradigm, SynTSBench, that systematically assesses fundamental modeling capabilities of time series forecasting models through programmable feature configuration. Our framework isolates confounding factors and establishes an interpretable evaluation system with three core analytical dimensions: (1) temporal feature decomposition and capability mapping, which enables systematic evaluation of model capacities to learn specific pattern types; (2) robustness analysis under data irregularities, which quantifies noise tolerance thresholds and anomaly recovery capabilities; and (3) theoretical optimum benchmarking, which establishes performance boundaries for each pattern type-enabling direct comparison between model predictions and mathematical optima. Our experiments show that current deep learning models do not universally approach optimal baselines across all types of temporal features.The code is available at https://github.com/TanQitai/SynTSBench

时序建模评估框架合成数据鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。