arXiv:2601.06100cs.LGcs.CL2026-01被引 1

用卡尔曼滤波解释大模型推理时学习,揭示不确定性的动态作用。

Filtering Beats Fine Tuning: A Bayesian Kalman View of In Context Learning in LLMs

  • 将上下文学习建模为在线贝叶斯状态估计,用线性状态空间描述低维适配状态。
  • 发现后验方差快速收缩是学习驱动力,先于均值收敛,且可证明指数级收缩速率。
  • 适用于理解提示有效性、参数高效微调与测试时学习,尤其适合研究不确定性机制者。

我们提出一个理论驱动的框架,将大语言模型(LLMs)推理时的适应性视为在线贝叶斯状态估计。不同于将快速适应视为隐式优化或元学习,我们将其建模为由线性化状态空间模型控制的低维潜在适配状态的序贯推断。在高斯假设下,适配遵循闭式更新的卡尔曼递归,同时更新后验均值与协方差。该视角将认知不确定性提升为显式动态变量。我们证明推理时的学习由协方差坍缩驱动,即信息性标记引发的后验不确定性快速收缩,通常早于后验均值收敛。基于标记级雅可比的可观性条件,我们建立了贝叶斯滤波的稳定性,证明了指数级协方差收缩率,并推导出均方误差界。梯度下降、自然梯度及元学习更新被视为滤波动态在无噪声极限下的退化解,表明基于优化的适配只是贝叶斯推断的退化近似。该理论统一解释了上下文学习、参数高效适配与测试时学习,提供稳定性与样本效率的明确保证,通过信息累积解释提示有效性,并阐明现有方法缺失的不确定性动态。少量示例实验验证了理论的定性预测。

原文摘要 · Abstract (English)

We present a theory-first framework that interprets inference-time adaptation in large language models (LLMs) as online Bayesian state estimation. Rather than modeling rapid adaptation as implicit optimization or meta-learning, we formulate task- and context-specific learning as the sequential inference of a low-dimensional latent adaptation state governed by a linearized state-space model. Under Gaussian assumptions, adaptation follows a Kalman recursion with closed-form updates for both the posterior mean and covariance. This perspective elevates epistemic uncertainty to an explicit dynamical variable. We show that inference-time learning is driven by covariance collapse, i.e., rapid contraction of posterior uncertainty induced by informative tokens, which typically precedes convergence of the posterior mean. Using observability conditions on token-level Jacobians, we establish stability of the Bayesian filter, prove exponential covariance contraction rates, and derive mean-square error bounds. Gradient descent, natural-gradient methods, and meta-learning updates arise as singular, noise-free limits of the filtering dynamics, positioning optimization-based adaptation as a degenerate approximation of Bayesian inference. The resulting theory provides a unified probabilistic account of in-context learning, parameter-efficient adaptation, and test-time learning without parameter updates. It yields explicit guarantees on stability and sample efficiency, offers a principled interpretation of prompt informativeness via information accumulation, and clarifies the role of uncertainty dynamics absent from existing accounts. Minimal illustrative experiments corroborate the qualitative predictions of the theory.

大模型贝叶斯推断上下文学习卡尔曼滤波

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。