arXiv:2510.04102cs.LGcs.NA2025-10被引 2

揭示神经网络外推失效的根本原因,解释为何其无法像物理定律般可靠预测未知场景。

Why Cannot Neural Networks Master Extrapolation? Insights from Physical Laws

  • 提出衡量模型外推能力的关键属性,从理论上解释深度学习在域外预测中的性能下降。
  • 实验证明现有深度学习架构普遍缺乏该属性,导致长程预测远逊于简单基线。
  • 为下一代具备强外推能力的预测模型设计提供理论指引,适合关注时间序列建模的研究者。

受基础模型(FMs)在语言建模中取得显著成功启发,人们越来越关注开发用于时间序列预测的基础模型,因其对科学与工程具有变革潜力。尽管在短时预测任务中已取得重大进展,但外推或长程预测仍难以实现,当前模型表现甚至不及简单基线。这与物理定律所具备的强大外推能力形成鲜明对比,引发对神经网络结构与物理定律本质差异的思考。本文识别并形式化了一项决定统计学习模型在训练域外预测精度的核心属性,解释了深度学习模型在外推场景中性能退化的根源。结合理论分析与实证结果,我们展示了该属性对现有深度学习架构的影响。研究不仅阐明了外推差距的根本原因,还为设计能够掌握外推能力的下一代预测模型指明方向。

原文摘要 · Abstract (English)

Motivated by the remarkable success of Foundation Models (FMs) in language modeling, there has been growing interest in developing FMs for time series prediction, given the transformative power such models hold for science and engineering. This culminated in significant success of FMs in short-range forecasting settings. However, extrapolation or long-range forecasting remains elusive for FMs, which struggle to outperform even simple baselines. This contrasts with physical laws which have strong extrapolation properties, and raises the question of the fundamental difference between the structure of neural networks and physical laws. In this work, we identify and formalize a fundamental property characterizing the ability of statistical learning models to predict more accurately outside of their training domain, hence explaining performance deterioration for deep learning models in extrapolation settings. In addition to a theoretical analysis, we present empirical results showcasing the implications of this property on current deep learning architectures. Our results not only clarify the root causes of the extrapolation gap but also suggest directions for designing next-generation forecasting models capable of mastering extrapolation.

外推能力时间序列基础模型物理规律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。