arXiv:2505.18373cs.LGcs.AI2025-05被引 4

自回归预训练自然催生上下文学习,无需额外设计

Next-token pretraining implies in-context learning

  • 从词元预测出发,用信息论框架解释上下文学习的必然性
  • 实验验证了训练损失相变与幂律缩放等典型现象
  • 揭示模型推理时学习能力与预训练任务的数学关联

我们主张,上下文学习(ICL)是标准自监督词元预测预训练的可预测结果,而非奇特的涌现属性。本文通过聚焦分布内ICL,阐明模型在序列数据上训练时必然适应上下文,尤其当数据源非遍历性时。我们提出的信息论框架可精确预测此类ICL动态(即上下文依赖的损失降低)。通过不同相关结构的合成数据集实验,验证了诱导头形成时训练损失的相变及上下文损失的幂律缩放等特征现象。进一步表明,模型在任一任务上的上下文性能,与预训练中经历的任务集合存在数学耦合,为推理时学习提供了基于架构和模态无关原理的根本解释。

原文摘要 · Abstract (English)

We argue that in-context learning (ICL) predictably arises from standard self-supervised next-token pretraining, rather than being an exotic emergent property. This work establishes the foundational principles of this emergence by focusing on in-distribution ICL, demonstrating how models necessarily adapt to context when trained on token sequences, especially from non-ergodic sources. Our information-theoretic framework precisely predicts these in-distribution ICL dynamics (i.e., context-dependent loss reduction). We verify this with experiments using synthetic datasets of differing types of correlational structure, reproducing characteristic phenomena like phase transitions in training loss for induction head formation and power-law scaling of in-context loss. We further show that a model's in-context performance on any task is mathematically coupled to the ensemble of tasks seen in pretraining, offering a fundamental explanation, grounded in architecture- and modality-independent principles, for such inference-time learning.

上下文学习自回归信息论预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。