arXiv:2512.22702cs.LG2025-12

现有基准测试无法识别模型性能差异根源,阻碍时序预测进步

Position: Current Benchmarking Hinders Real Progress in Deep Learning for Time Series Forecasting

  • 对比模型时忽略关键设计差异,如全局性与局部性
  • 设计维度差异的影响可能超过使用特定建模层
  • 适合关注模型可复现性与公平比较的研究者

深度学习模型在时序应用中日益流行,但新架构层出不穷,实证结果常互相矛盾,难以判断何种设计选择真正提升性能。本文认为,当前基准测试方法无法识别性能差异的根本原因,导致领域进展缓慢。特别地,模型比较中常被忽视的关键设计维度(如全局性与局部性)会从根本上改变预测方法类别,并显著影响实验结果。我们发现,这些常被视为实现细节的差异,其影响甚至超过采用特定序列建模层。为此,我们建议重新审视基准测试实践,聚焦预测问题的基础特性。作为具体举措,提出一个辅助预测模型卡片模板,包含一组字段,用于基于关键设计选择表征现有及新架构。

原文摘要 · Abstract (English)

Deep learning models have grown popular in time series applications. However, the large quantity of newly proposed architectures and the often contradictory empirical results make it difficult to assess which design choice and model component drives performance. In this position paper, we argue that current benchmarking practices fail to identify the factors responsible for performance differences, thus slowing down progress in the field. In particular, differences in crucial design dimensions are overlooked when comparing architectures, ultimately leading to inconsistent outcomes. To support our position, we show that such differences-often treated as mere implementation details-can have a greater impact than adopting specific sequence modeling layers. We discuss how overlooked aspects (such as globality and locality) can (1) fundamentally change the class of the forecasting method and (2) drastically affect empirical results. Our findings suggest rethinking our benchmarking practices and focusing on the foundational aspects of the forecasting problem when designing and comparing architectures. As a concrete step, we propose an auxiliary forecasting model card, i.e., a template with a set of fields to characterize existing and new forecasting architectures based on key design choices.

时序预测基准测试模型设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。