揭示大模型上下文学习源于序列与主题的泛化能力
Towards Auto-Regressive Next-Token Prediction: In-Context Learning Emerges from Generalization
- 用自回归预测构建更贴近真实语言训练的上下文学习框架
- 提出依赖数据、主题和优化过程的泛化边界理论
- 首次从泛化角度解释上下文学习如何涌现,适合研究模型机制者
大型语言模型在上下文学习(ICL)方面表现卓越,但现有理论分析存在两大局限:(a) 仅限独立同分布(i.i.d.)设定,多基于随机输入-标签对构建提示,脱离真实语言中提示词间的相互依赖;(b) 缺乏涌现机制解释,多数研究仅从隐式优化角度说明ICL功能,未阐明其生成过程及预训练阶段的影响。本文通过采用更贴近实际训练的自回归下一词预测(AR-NTP)范式,强调提示词之间的序列依赖性——即每个词的预测依赖于前序完整序列。同时,我们建立系统化的预训练与ICL理论框架,揭示序列与主题的分层结构及双重期望机制。最终,推导出依赖数据、主题和优化过程的PAC-Bayesian泛化界,表明ICL源于对序列与主题的泛化能力。实验在数值线性动态系统、合成GINC及真实语言数据集上验证了该理论。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated remarkable in-context learning (ICL) abilities. However, existing theoretical analysis of ICL primarily exhibits two limitations: (a) Limited i.i.d. Setting. Most studies focus on supervised function learning tasks where prompts are constructed with i.i.d. input-label pairs. This i.i.d. assumption diverges significantly from real language learning scenarios where prompt tokens are interdependent. (b) Lack of Emergence Explanation. Most literature answers what ICL does from an implicit optimization perspective but falls short in elucidating how ICL emerges and the impact of pre-training phase on ICL. In our paper, to extend (a), we adopt a more practical paradigm, auto-regressive next-token prediction (AR-NTP), which closely aligns with the actual training of language models. Specifically, within AR-NTP, we emphasize prompt token-dependency, which involves predicting each subsequent token based on the preceding sequence. To address (b), we formalize a systematic pre-training and ICL framework, highlighting the layer-wise structure of sequences and topics, alongside a two-level expectation. In conclusion, we present data-dependent, topic-dependent and optimization-dependent PAC-Bayesian generalization bounds for pre-trained LLMs, investigating that ICL emerges from the generalization of sequences and topics. Our theory is supported by experiments on numerical linear dynamic systems, synthetic GINC and real-world language datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。