arXiv:2410.01565cs.LGstat.ML2024-10被引 6

用贝叶斯后验解释上下文学习,揭示模型泛化新机制

Bayes' Power for Explaining In-Context Learning Generalizations

  • 将神经网络视为数据生成过程的后验近似,而非最大似然估计
  • 实验证明模型通过组合训练数据知识实现惊人泛化能力
  • 适合研究大模型推理机制与泛化理论的研究者阅读

传统上,神经网络训练被视为最大似然估计(MLE)的近似。这一观点源于小数据集多轮训练的时代,当时性能受限于数据量;但在大规模单轮训练(如自监督语言模型)时代,性能转为受算力约束,数据则充足。随着模型能力增强,上下文学习(ICL)——即基于上下文在一次前向传播中完成学习——成为主流范式。本文提出,在此背景下,更合理的视角是将神经网络行为理解为对真实后验的近似,该后验由数据生成过程定义。我们展示了这一解释在解释ICL及其预测未见任务泛化方面的强大能力。实验表明,模型通过有效整合训练数据中的知识,成为稳健的上下文学习者。所有发现均可通过精确后验解释。最后,我们揭示了后验固有的泛化限制,以及神经网络在逼近这些后验时的局限性。

原文摘要 · Abstract (English)

Traditionally, neural network training has been primarily viewed as an approximation of maximum likelihood estimation (MLE). This interpretation originated in a time when training for multiple epochs on small datasets was common and performance was data bound; but it falls short in the era of large-scale single-epoch trainings ushered in by large self-supervised setups, like language models. In this new setup, performance is compute-bound, but data is readily available. As models became more powerful, in-context learning (ICL), i.e., learning in a single forward-pass based on the context, emerged as one of the dominant paradigms. In this paper, we argue that a more useful interpretation of neural network behavior in this era is as an approximation of the true posterior, as defined by the data-generating process. We demonstrate this interpretations' power for ICL and its usefulness to predict generalizations to previously unseen tasks. We show how models become robust in-context learners by effectively composing knowledge from their training data. We illustrate this with experiments that reveal surprising generalizations, all explicable through the exact posterior. Finally, we show the inherent constraints of the generalization capabilities of posteriors and the limitations of neural networks in approximating these posteriors.

贝叶斯推理上下文学习泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。