arXiv:2605.03780cs.LGcs.CL2026-05

揭示Transformer模型任务推理的两种模式及其几何原理

Task Vector Geometry Underlies Dual Modes of Task Inference in Transformers

论文配图:Task Vector Geometry Underlies Dual Modes of Task Inference in Transformers
图 1 · 摘自论文原文
  • 通过合成数据训练小模型,研究任务向量的几何结构
  • 同一模型内可同时实现已见任务识别与新任务适应
  • 任务向量空间正交区域支持对未知任务的外推泛化

Transformer在从上下文中推断潜在任务时具备两种推理模式:识别训练中见过的任务,以及适应新任务。近期可解释性研究发现,中间层表征中存在特定方向的任务向量,可引导模型行为。然而,缺乏严格的理论基础阻碍了内部表征与外部行为之间的联系:现有工作无法解释任务向量几何如何由训练分布决定,以及何种几何结构支持分布外(OOD)泛化。本文在受控的合成设置中,从零训练小型Transformer模型,学习潜在任务序列分布,实现严谨的数学刻画。结果表明,两种推理模式可在同一模型中共存。分布内行为由贝叶斯任务检索驱动,通过学习到的任务向量的凸组合实现;而分布外行为则源于外推式任务学习,其表征位于几乎与任务向量子空间正交的子空间中。整体表明,任务向量几何、训练分布与泛化行为之间存在紧密关联。

原文摘要 · Abstract (English)

Transformers are effective at inferring the latent task from context via two inference modes: recognizing a task seen during training, and adapting to a novel one. Recent interpretability studies have identified from middle-layer representations task-specific directions, or task vectors, that steer model behavior. However, a lack of rigorous foundations hinders connecting internal representations to external model behavior: existing work fails to explain how task-vector geometry is shaped by the training distribution, and what geometry enables out-of-distribution (OOD) generalization. In this paper, we study these questions in a controlled synthetic setting by training small transformers from scratch on latent-task sequence distributions, which allows a principled mathematical characterization. We show that two inference modes can coexist within a single model. In-distribution behavior is governed by Bayesian task retrieval, implemented internally through convex combinations of learned task vectors. OOD behavior, by contrast, arises through extrapolative task learning, whose representations occupy a subspace nearly orthogonal to the task-vector subspace. Taken together, our results suggest that task-vector geometry, training distributions, and generalization behaviors are closely related.

Transformer任务推理可解释性几何结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。