arXiv:2601.22364cs.CLcs.AI2026-01被引 7

研究发现大模型在不同任务中会动态调整表示结构,有的变直提升预测,有的则不。

Context Structure Reshapes the Representational Geometry of Language Models

  • 通过测量上下文内表示轨迹的直线度,分析模型如何随任务变化调整内部表征。
  • 连续预测任务中上下文越长轨迹越直,预测性能越好;结构化任务中仅部分阶段出现直线化。
  • 揭示了大模型在上下文学习中并非统一机制,而是根据任务灵活切换策略。

大型语言模型(LLMs)在深层中将输入序列的表示组织为更直的神经轨迹,这一现象被认为有助于通过线性外推进行下一个词预测。模型还能适应多种任务并在上下文中学习新结构,近期研究显示这种上下文学习(ICL)会反映在表示变化中。本文将这两类研究结合,探索在上下文学习过程中是否会发生表示直线化。我们在Gemma 2模型上对多样化的上下文任务进行了测量,发现模型表示的变化存在两极分化:在持续预测任务(如自然语言、网格世界遍历)中,随着上下文增长,神经序列轨迹变得更直,且与预测性能提升相关;而在结构化预测任务(如少样本任务)中,直线化不一致——仅在具有明确结构的阶段(如重复模板)出现,其余阶段则消失。结果表明,上下文学习并非单一过程。我们提出,大模型如同瑞士军刀,根据任务结构动态选择策略,仅部分策略会产生表示直线化。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have been shown to organize the representations of input sequences into straighter neural trajectories in their deep layers, which has been hypothesized to facilitate next-token prediction via linear extrapolation. Language models can also adapt to diverse tasks and learn new structure in context, and recent work has shown that this in-context learning (ICL) can be reflected in representational changes. Here we bring these two lines of research together to explore whether representation straightening occurs \emph{within} a context during ICL. We measure representational straightening in Gemma 2 models across a diverse set of in-context tasks, and uncover a dichotomy in how LLMs' representations change in context. In continual prediction settings (e.g., natural language, grid world traversal tasks) we observe that increasing context increases the straightness of neural sequence trajectories, which is correlated with improvement in model prediction. Conversely, in structured prediction settings (e.g., few-shot tasks), straightening is inconsistent -- it is only present in phases of the task with explicit structure (e.g., repeating a template), but vanishes elsewhere. These results suggest that ICL is not a monolithic process. Instead, we propose that LLMs function like a Swiss Army knife: depending on task structure, the LLM dynamically selects between strategies, only some of which yield representational straightening.

语言模型表示几何上下文学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。