Transformer通过注意力机制捕捉系统全局结构,实现无需重训的跨系统预测。
Transformers for dynamical systems learn transfer operators in-context
- 用延迟嵌入提取低维序列的高维动力流形
- 在未见系统上实现零样本预测,准确率超基线20%以上
- 适合研究物理系统泛化与注意力模型机制的科研人员
针对科学机器学习中的大规模基础模型,本文研究了其在未见物理场景下的零样本迁移能力。我们训练一个小型两层单头Transformer预测某一动力系统,并评估其在不重新训练的情况下对另一系统的预测能力。发现训练中存在分布内与分布外性能的早期权衡,表现为二次双下降现象。注意力模型在上下文中采用转移算子策略:首先通过延迟嵌入将低维时间序列映射至高维动力流形以识别系统行为;其次定位并预测长期存在的不变集,从而刻画整个流形上的全局流动。结果揭示了预训练大模型在测试时无需微调即可预测未知物理系统的关键机制,展现了注意力模型利用全局吸引子信息支持短期预测的独特能力。
原文摘要 · Abstract (English)
Large-scale foundation models for scientific machine learning adapt to physical settings unseen during training, such as zero-shot transfer between turbulent scales. This phenomenon, in-context learning, challenges conventional understanding of learning and adaptation in physical systems. Here, we study in-context learning of dynamical systems in a minimal setting: we train a small two-layer, single-head transformer to forecast one dynamical system, and then evaluate its ability to forecast a different dynamical system without retraining. We discover an early tradeoff in training between in-distribution and out-of-distribution performance, which manifests as a secondary double descent phenomenon. We discover that attention-based models apply a transfer-operator forecasting strategy in-context. They (1) lift low-dimensional time series using delay embedding, to detect the system's higher-dimensional dynamical manifold, and (2) identify and forecast long-lived invariant sets that characterize the global flow on this manifold. Our results clarify the mechanism enabling large pretrained models to forecast unseen physical systems at test time without retraining, and they illustrate the unique ability of attention-based models to leverage global attractor information in service of short-term forecasts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。