arXiv:2505.11349cs.LGnlin.CD2025-05被引 13

简单复制上下文竟比复杂模型更准,揭示科学机器学习中的隐藏漏洞

Context parroting: A simple but tough-to-beat baseline for foundation models in scientific machine learning

  • 用直接复制输入片段的方式做预测,效果优于主流时间序列大模型
  • 在混沌系统、湍流等多类动态系统上,该方法准确率更高且成本极低
  • 揭示了当前模型依赖记忆而非理解,适合研究模型泛化与内训机制

近期时间序列基础模型在预测物理系统方面表现出强大能力,包括仅凭短轨迹上下文进行零样本预测,无需了解底层物理规律。本文发现,这些模型往往通过简单的‘上下文鹦鹉学舌’策略进行预测;当不采用该策略时,则普遍出现收敛到均值等共性失败模式。一个仅直接复制上下文的朴素模型,在多种动力系统(如低维混沌、湍流、耦合振子、心电图)上的预测表现优于领先的时间序列基础模型,且计算成本仅为后者的极小部分。我们指出上下文鹦鹉学舌与归纳头(induction heads)的相似性,解释了为何大语言模型可被重用于时间序列预测。从动力系统视角出发,本文还将预测精度与上下文长度的缩放关系联系到混沌吸引子的分形维数,为先前观测到的上下文内神经缩放律提供了新见解。该发现揭示了现有模型的性能差距与失效模式,有助于未来基础模型的设计及超越鹦鹉学舌的上下文学习策略探索。

原文摘要 · Abstract (English)

Recent time-series foundation models exhibit strong abilities to predict physical systems. These abilities include zero-shot forecasting, in which a model forecasts future states of a system given only a short trajectory as context, without knowledge of the underlying physics. Here, we show that foundation models often forecast through a simple parroting strategy, and when they are not parroting they exhibit some shared failure modes such as converging to the mean. As a result, a naive context parroting model that copies directly from the context scores higher than leading time-series foundation models on predicting a diverse range of dynamical systems, including low-dimensional chaos, turbulence, coupled oscillators, and electrocardiograms, at a tiny fraction of the computational cost. We draw a parallel between context parroting and induction heads, which explains recent works showing that large language models can often be repurposed for time series forecasting. Our dynamical systems perspective also ties the scaling between forecast accuracy and context length to the fractal dimension of the underlying chaotic attractor, providing insight into previously observed in-context neural scaling laws. By revealing the performance gaps and failure modes of current time-series foundation models, context parroting can guide the design of future foundation models and help identify in-context learning strategies beyond parroting.

时间序列基础模型零样本预测混沌系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。