用贝叶斯视角揭示大模型延迟泛化背后的不确定性机制
A Bayesian Perspective on the Role of Epistemic Uncertainty for Delayed Generalization in In-Context Learning

- 通过贝叶斯方法分析模型在上下文学习中的预测不确定性变化
- 发现模型突现泛化时认知不确定性急剧下降,可作为无标签诊断指标
- 理论证明延迟泛化与不确定性峰值由同一谱机制驱动,适用于研究模型训练动态
上下文学习使Transformer在推理时仅凭少量示例即可适应新任务,而‘突现’现象则表明这种泛化可能在长期训练后突然出现。本文从贝叶斯视角研究上下文学习中的任务泛化与突现现象,探究从记忆到泛化的延迟转变由何驱动。具体考察模块算术任务中,Transformer仅从上下文示例推断隐含线性函数的性能,并结合近似贝叶斯技术估计后验分布,分析不确定性在训练过程中的演变规律,以及在任务多样性、上下文长度和噪声变化下的行为。结果发现,当模型发生突现时,认知不确定性会急剧塌缩,使其成为衡量泛化的无标签实用指标。此外,通过简化贝叶斯线性模型提供理论支持,表明延迟泛化与不确定性峰值均源于相同的谱机制,揭示了突现时间与不确定性动态之间的内在联系。
原文摘要 · Abstract (English)
In-context learning enables transformers to adapt to new tasks from a few examples at inference time, while grokking highlights that this generalization can emerge abruptly only after prolonged training. We study task generalization and grokking in in-context learning using a Bayesian perspective, asking what enables the delayed transition from memorization to generalization. Concretely, we consider modular arithmetic tasks in which a transformer must infer a latent linear function solely from in-context examples and analyze how predictive uncertainty evolves during training. We combine approximate Bayesian techniques to estimate the posterior distribution and we study how uncertainty behaves across training and under changes in task diversity, context length, and context noise. We find that epistemic uncertainty collapses sharply when the model groks, making uncertainty a practical label-free diagnostic of generalization in transformers. Additionally, we provide theoretical support with a simplified Bayesian linear model, showing that asymptotically both delayed generalization and uncertainty peaks arise from the same underlying spectral mechanism, which links grokking time to uncertainty dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。