arXiv:2602.17744cs.LGcs.CL2026-02

提出贝叶斯最优序列预测新范式,解释上下文学习为何高效。

Bayesian Optimality of In-Context Learning with Selective State Spaces

  • 将上下文学习建模为隐序列任务的元学习,用选择性状态空间模型实现贝叶斯最优预测。
  • 在线性高斯状态空间模型下,渐近收敛至后验预测均值,优于经验风险最小化方法。
  • 适合关注模型统计效率与架构设计原理的研究者,尤其适用于结构化噪声场景。

我们提出贝叶斯最优序列预测作为理解上下文学习(ICL)的新原则。不同于将Transformer视为隐式梯度下降的解释,我们将其形式化为对隐序列任务的元学习。对于由线性高斯状态空间模型(LG-SSMs)支配的任务,我们证明元训练的选择性状态空间模型渐近地实现了贝叶斯最优预测器,收敛至后验预测均值。进一步,我们建立了与梯度下降的统计分离:构造了具有时间相关噪声的任务,其中最优贝叶斯预测器严格优于任何经验风险最小化(ERM)估计器。由于Transformer可视为执行隐式ERM,这表明选择性状态空间模型因更高的统计效率而达到更低的渐近风险。在合成LG-SSM任务和字符级马尔可夫基准上的实验验证:选择性状态空间模型更快收敛至贝叶斯最优风险,在结构化噪声设置中更长上下文下样本效率更高,且对潜在状态的追踪比线性Transformer更鲁棒。这一框架将ICL从‘隐式优化’重新定义为‘最优推断’,解释了选择性状态空间模型的效率,并为架构设计提供了原则性基础。

原文摘要 · Abstract (English)

We propose Bayesian optimal sequential prediction as a new principle for understanding in-context learning (ICL). Unlike interpretations framing Transformers as performing implicit gradient descent, we formalize ICL as meta-learning over latent sequence tasks. For tasks governed by Linear Gaussian State Space Models (LG-SSMs), we prove a meta-trained selective SSM asymptotically implements the Bayes-optimal predictor, converging to the posterior predictive mean. We further establish a statistical separation from gradient descent, constructing tasks with temporally correlated noise where the optimal Bayesian predictor strictly outperforms any empirical risk minimization (ERM) estimator. Since Transformers can be seen as performing implicit ERM, this demonstrates selective SSMs achieve lower asymptotic risk due to superior statistical efficiency. Experiments on synthetic LG-SSM tasks and a character-level Markov benchmark confirm selective SSMs converge faster to Bayes-optimal risk, show superior sample efficiency with longer contexts in structured-noise settings, and track latent states more robustly than linear Transformers. This reframes ICL from "implicit optimization" to "optimal inference," explaining the efficiency of selective SSMs and offering a principled basis for architecture design.

上下文学习贝叶斯推断状态空间模型元学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。