arXiv:2608.28059cond-mat.dis-nncs.LG2026-08

用统计物理解释线性上下文学习中的奇异误差现象

Landau theory of quenched criticality in linear in-context learning

论文配图:Landau theory of quenched criticality in linear in-context learning
图 1 · 摘自论文原文
  • 将上下文学习误差奇点建模为非平衡系统的临界现象
  • 发现样本间参数波动是误差发散的微观根源,临界点在τ=1
  • 理论预测与数值模拟高度吻合,适用于理解模型泛化

上下文学习(ICL)使预训练模型在不更新参数的情况下,仅通过提示中的示例推断新任务。在线性ICL模型中,当预训练样本数接近可学习参数数量时,预测误差出现双下降奇点。我们将这一插值奇点视为一种淬火无序系统的临界现象。通过对比遍历与淬火描述,识别出学习参数的样本间关联波动是误差奇点的微观起源。通过积分腔自洽方程,构建了重整化岭参数ξ的朗道势。ξ扮演有序参数角色,原始岭参数λ为其共轭磁场,归一化样本复杂度τ相当于温度,双下降奇点出现在临界温度τ_c=1。朗道顺从度恰好对应于预测误差中波动贡献的发散量。有序参数与无岭极限下经验松弛矩阵零特征值占比密切相关,定义了学习动力学中的平坦方向。该朗道理论具有通用立方形式,临界指数为(β_cr,δ_cr,γ_cr)=(1,2,1)。大上下文区域出现类赝隙区,有序参数被抑制。朗道理论的预测经由原学习问题的数值求解独立验证,定量吻合良好。本研究为线性上下文学习中的插值临界性提供了坚实的统计物理理解。

原文摘要 · Abstract (English)

In-context learning (ICL) allows a pretrained model to infer a new task from examples supplied in its prompt without updating its parameters. In linear models of ICL, the prediction error develops a double-descent singularity when the number of pretraining samples becomes comparable to the number of learnable parameters. We formulate this interpolation singularity as a critical phenomenon of a quenched disordered system. By comparing annealed and quenched descriptions of the same linear ICL model, we identify the connected sample-to-sample fluctuations of the learned parameters as the microscopic origin of the singular error. A Landau potential is constructed by integrating the cavity self-consistency equation for the renormalized ridge parameter $ξ$. The role of (magnetization) order parameter is played by $ξ$, while the bare ridge parameter $λ$ becomes its conjugate magnetic field. The normalized sample complexity $τ$ acts as a temperature and the double-descent singularity occurs at the critical temperature $τ_c =1$. The Landau susceptibility is precisely the quantity that diverges in the fluctuation contribution to the prediction error. The order parameter is closely related to the fraction of zero eigenvalues of the empirical relaxation matrix in the ridgeless limit, which define flat directions in the learning dynamics. The Landau theory is generically cubic in the order parameter with critical exponents $(β_{\rm cr},δ_{\rm cr},γ_{\rm cr})=(1,2,1)$. In the large-context regime, there appears a pseudogap-like regime characterized by suppressed order parameter. Predictions of the Landau theory are independently confirmed from numerical solutions of the original learning problem with good quantitative agreement. Our results pave the way for solid statistical-physics understanding of the interpolation criticality in linear in-context learning.

上下文学习统计物理临界现象双下降

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。