LLM通过重构内部表征几何结构实现上下文学习,不同任务难易取决于表征可重排程度。
Large language models reorganize representational geometry during in-context learning

- 构造基于模型自身表征投影的线性分类任务,研究其可学习性
- 成功学习伴随表征几何重排,提升任务相关可分性,但激活特定轴无效
- 行为符合原型算法,表征在上下文中被重新组织以适应任务
大型语言模型(LLMs)无需参数更新即可灵活适应新任务,这种能力称为上下文学习(ICL)。然而,为何某些任务易学而另一些难学仍不明确。本文研究了LLMs能否任意调整其表征以解决一个简单的线性分类任务。具体地,构建了一组二分类任务,标签由模型自身表征投影到不同轴上定义。尽管所有任务在构造上均为线性可分,但其在上下文学习中的可学习性随投影轴系统性变化。发现成功进行ICL时,内部表征的几何结构会重新组织,从而增强任务相关的可分性。因果干预中,仅增强任务定义轴上的神经活动不足以提升行为表现或引发表征重排。此外,模型行为最符合基于重新组织表征的原型类算法。这些结果为LLM的ICL提供了几何解释,表明训练获得的表征限制了上下文学习可利用的潜力。
原文摘要 · Abstract (English)
Large language models (LLMs) show remarkable flexibility in adapting to novel tasks without parameter updates, a capacity known as in-context learning (ICL). Prior work has sought to understand ICL by studying the circuits, algorithms, and representations that support it. Yet why some ICL tasks are easy to solve while others are difficult remains unresolved. In this paper, we ask whether LLMs can adapt their representations arbitrarily to solve a simple linear classification task. Specifically, we construct a family of binary classification tasks in which labels are defined by projecting LLMs' own representations onto different axes. Surprisingly, although all tasks are linearly separable by construction, their in-context learnability varies systematically across axes. We find that successful ICL is accompanied by a geometric reorganization of internal representations that increases task-relevant separability. Causal interventions that amplify neural activity along the axis defining the task are insufficient to improve behavioral performance or induce this representational reorganization. We also show that LLM behavior is best described by a prototype-like algorithm operating on representations that are themselves reorganized in context to adapt to the task. Together, these findings offer a geometric account of ICL in LLMs, showing that representations acquired through training constrain what can be exploited through in-context learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。