大模型在无物理训练下自发学习能量等物理概念,解释其推理机制。
Uncovering Spontaneous Physics Representations in In-Context Learning
- 通过上下文预测物理动态,逐步形成能量等物理量的内部表征。
- 上下文越长,模型预测越准,且能量相关激活增强。
- 能量信号影响预测结果,说明模型真正理解物理规律而非简单记忆。
上下文学习(ICL)使大语言模型(LLMs)仅凭提示即可解决新任务,涵盖广泛领域,但其内在机制仍不明确。物理系统因其结构化动态和可控制数据,成为研究该问题的理想实验场。本文聚焦于物理推理中的ICL能力,以动力学预测为代理任务。结果表明,随着上下文历史长度增加,LLMs对物理动态的预测准确率提升。分析模型残差流发现,其内部激活与能量等关键物理量存在相关性,且该相关性随上下文长度逐渐增强,表明模型在无物理专项监督下自发形成符合物理概念的表示。进一步引入逐层梯度归因分析发现,与能量相关性更强的残差方向在数值预测中获得更高贡献,而与位移等直接观测量相关的特征则无此模式,说明能量信号并非简单复制输入。本研究将ICL分析扩展至结构化物理动态,揭示了模型在上下文中组织物理结构的机制。
原文摘要 · Abstract (English)
In-context learning (ICL) lets large language models (LLMs) solve new tasks from prompts alone, across an ever-widening range of domains, yet the mechanisms underlying this ability remain poorly understood. Physical systems offer a controlled testbed for this question as they provide experimentally controllable data with structured dynamics grounded in fundamental principles. Here we study the ICL ability of LLMs, focusing on physical reasoning. Using dynamics forecasting as a proxy task, we first show that LLMs forecast physical dynamics in context, with accuracy improving as more history is provided. Analyzing the model's residual stream reveals internal activations that correlate with key physical quantities such as energy. These correlations strengthen gradually with context length, indicating that LLMs spontaneously form representations aligned with physical concepts without any physics-specific supervision. To assess whether these representations contribute to the model's predictions, we introduce a layer-wise gradient-based attribution analysis. We find that, residual directions more strongly correlated with energy also receive greater attribution to numerical predictions. This pattern is not observed for features correlated with directly observed quantities such as displacement, suggesting that the energy-related signal is not merely numerical information copied from the input. Our results broaden ICL analysis to structured physical dynamics and give a mechanistic account of how LLMs organize physical structure in context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。