文本训练的大模型无需微调即可零样本推演偏微分方程动态,揭示三阶段上下文学习机制。
Text-Trained LLMs Can Zero-Shot Extrapolate PDE Dynamics, Revealing a Three-Stage In-Context Learning Mechanism
- 利用上下文学习直接推演离散化偏微分方程解,不需微调或提示词
- 预测误差随时间步数代数增长,与经典数值求解器类似
- 发现三阶段学习过程:模仿语法→探索高熵→精准数值预测
大型语言模型(LLMs)在多种任务中展现出涌现的上下文学习(ICL)能力,包括零样本时间序列预测。我们发现,仅通过文本训练的基础模型可在无需微调或自然语言提示的情况下,准确外推离散化偏微分方程(PDE)解的时空动态。预测精度随时间上下文长度增加而提升,但在更细的空间离散化下下降。在多步滚动预测中,模型递归地预测未来空间状态,误差随时间范围代数增长,类似于经典有限差分求解器中的全局误差累积。我们将其解释为一种上下文神经尺度律,即预测质量可预测地随上下文长度和输出长度变化。为深入理解模型如何内部处理PDE解以实现准确滚动预测,我们分析了标记级输出分布,揭示出一致的三阶段ICL演化过程:初始为句法模式模仿,过渡至探索性高熵阶段,最终进入自信且数值基础牢固的预测阶段。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated emergent in-context learning (ICL) capabilities across a range of tasks, including zero-shot time-series forecasting. We show that text-trained foundation models can accurately extrapolate spatiotemporal dynamics from discretized partial differential equation (PDE) solutions without fine-tuning or natural language prompting. Predictive accuracy improves with longer temporal contexts but degrades at finer spatial discretizations. In multi-step rollouts, where the model recursively predicts future spatial states over multiple time steps, errors grow algebraically with the time horizon, reminiscent of global error accumulation in classical finite-difference solvers. We interpret these trends as in-context neural scaling laws, where prediction quality varies predictably with both context length and output length. To better understand how LLMs are able to internally process PDE solutions so as to accurately roll them out, we analyze token-level output distributions and uncover a consistent three-stage ICL progression: beginning with syntactic pattern imitation, transitioning through an exploratory high-entropy phase, and culminating in confident, numerically grounded predictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。