揭示大模型上下文学习中低频特征的内在偏置机制
Provable Low-Frequency Bias of In-Context Learning of Representations
- 提出双收敛框架,解释隐藏表示如何在上下文和层间同步收敛
- 证明表示趋向平滑低频结构,总能量衰减但不消失
- 解释几何扭曲现象并预测对高频噪声的鲁棒性,适合理论研究者
上下文学习(ICL)使大语言模型仅通过输入序列即可获得新行为,无需参数更新。近期研究表明,ICL可通过内化提示数据生成过程(DGP)的结构来超越预训练阶段的原始语义。然而其机制仍不明确。本文首次提出统一的双收敛框架,揭示隐藏表示在上下文与层间同时收敛的规律,导致隐含的平滑(低频)表示偏好。我们从理论上证明该现象,并通过实验证实。理论可解释多个开放性观察,包括为何学习表示呈现全局结构但局部扭曲的几何特性,以及为何总能量衰减却不消失。此外,理论预测ICL对高频噪声具有内在鲁棒性,实验予以验证。这些结果为理解ICL机制提供新视角与理论基础,有望推广至更广泛的数据分布与场景。
原文摘要 · Abstract (English)
In-context learning (ICL) enables large language models (LLMs) to acquire new behaviors from the input sequence alone without any parameter updates. Recent studies have shown that ICL can surpass the original meaning learned in pretraining stage through internalizing the structure the data-generating process (DGP) of the prompt into the hidden representations. However, the mechanisms by which LLMs achieve this ability is left open. In this paper, we present the first rigorous explanation of such phenomena by introducing a unified framework of double convergence, where hidden representations converge both over context and across layers. This double convergence process leads to an implicit bias towards smooth (low-frequency) representations, which we prove analytically and verify empirically. Our theory explains several open empirical observations, including why learned representations exhibit globally structured but locally distorted geometry, and why their total energy decays without vanishing. Moreover, our theory predicts that ICL has an intrinsic robustness towards high-frequency noise, which we empirically confirm. These results provide new insights into the underlying mechanisms of ICL, and a theoretical foundation to study it that hopefully extends to more general data distributions and settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。