arXiv:2509.24169cs.CL2025-09被引 1

提出可直接训练的任务向量,提升模型推理性能并揭示其作用机制

Task Vectors, Learned Not Extracted: Performance Gains and Mechanistic Insight

  • 直接训练任务向量(LTVs),无需复杂提取过程
  • 低层通过注意力头的OV电路主导预测,少数关键头起决定作用
  • 早期向量旋转至相关子空间,后期主要放大幅度,具线性传播特性

大语言模型可通过上下文示例实现上下文学习(ICL)。近期研究认为这些示例被压缩为任务向量(TVs),即紧凑的任务表征。然而,现有方法通常通过复杂且不透明的方式从输出或隐藏状态中提取TVs,且很少阐明其计算机制。本文提出直接训练的学得任务向量(LTVs),在准确率上超越提取的TVs,且可在任意层、位置甚至不同ICL提示下有效使用。通过系统分析发现:底层上,TVs主要通过注意力头的输出-值(OV)电路影响预测,少数“关键头”最为重要;高层上,尽管存在Transformer非线性,但TV传播整体近似线性——早期向量被旋转至任务相关子空间以提升相关标签得分,后期则主要通过幅度缩放实现。LTVs不仅提供高效获取有效任务向量的方法,更揭示了ICL的机制基础。

原文摘要 · Abstract (English)

Large Language Models (LLMs) can perform new tasks from in-context demonstrations, a phenomenon known as in-context learning (ICL). Recent work suggests that these demonstrations are compressed into task vectors (TVs), compact task representations that LLMs exploit for predictions. However, prior studies typically extract TVs from model outputs or hidden states using cumbersome and opaque methods, and they rarely elucidate the mechanisms by which TVs influence computation. In this work, we address both limitations. First, we propose directly training Learned Task Vectors (LTVs), which surpass extracted TVs in accuracy and exhibit superior flexibility-acting effectively at arbitrary layers, positions, and even with ICL prompts. Second, through systematic analysis, we investigate the mechanistic role of TVs, showing that at the low level they steer predictions primarily through attention-head OV circuits, with a small subset of "key heads" most decisive. At a higher level, we find that despite Transformer nonlinearities, TV propagation is largely linear: early TVs are rotated toward task-relevant subspaces to improve logits of relevant labels, while later TVs are predominantly scaled in magnitude. Taken together, LTVs not only provide a practical approach for obtaining effective TVs but also offer a principled lens into the mechanistic foundations of ICL.

大模型任务向量机制解析上下文学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。