将上下文学习视为一种隐式知识蒸馏,解释其原理并提供理论支持。
Brewing Knowledge in Context: Distillation Perspectives on In-Context Learning
- 把上下文学习看作推理时的知识蒸馏过程,通过提示示例生成任务特定参考模型。
- 证明蒸馏权重偏差与提示和目标分布的MMD呈线性关系,理论解释了经验现象。
- 为提示工程和自动示例选择提供新视角,适合研究大模型推理机制的人阅读。
上下文学习(ICL)使大语言模型(LLMs)在无需更新权重的情况下解决新任务。尽管其在实践中表现良好,但其内在机制仍不清晰,限制了对它的解释、改进和可靠应用。本文提出一种新的理论视角,将ICL视为一种隐式的知识蒸馏(KD),其中提示示例引导模型在推理过程中形成任务特定的参考模型。在此框架下,我们推导出基于Rademacher复杂度的泛化界,并证明蒸馏权重的偏差随提示与目标分布之间的最大均值差异(MMD)线性增长。该理论框架解释了多个经验现象,并统一了先前基于梯度和分布的分析。据我们所知,这是首个将推理时注意力形式化为蒸馏过程的工作,为未来的提示工程和自动化示例选择提供了理论洞见。
原文摘要 · Abstract (English)
In-context learning (ICL) allows large language models (LLMs) to solve novel tasks without weight updates. Despite its empirical success, the mechanism behind ICL remains poorly understood, limiting our ability to interpret, improve, and reliably apply it. In this paper, we propose a new theoretical perspective that interprets ICL as an implicit form of knowledge distillation (KD), where prompt demonstrations guide the model to form a task-specific reference model during inference. Under this view, we derive a Rademacher complexity-based generalization bound and prove that the bias of the distilled weights grows linearly with the Maximum Mean Discrepancy (MMD) between the prompt and target distributions. This theoretical framework explains several empirical phenomena and unifies prior gradient-based and distributional analyses. To the best of our knowledge, this is the first to formalize inference-time attention as a distillation process, which provides theoretical insights for future prompt engineering and automated demonstration selection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。