arXiv:2606.07560cs.CLcs.LG2026-06

发现提示学习中注意力头分两类:写手和消音者,功能互补。

Function-Vector Heads Are Two Populations: Writers and Cancellers in In-Context Learning

  • 通过路径修补验证注意力头的符号方向,区分正向支持与反向抑制。
  • 消音头零删除使正确标签得分提升0.13~0.29纳特,准确率全数上升。
  • 该分工在六种任务、三类架构中稳定存在,适于理解模型内部机制。

函数向量(FV)分析通常根据注意力头对上下文任务的因果贡献大小识别其重要性,但忽略效应方向。本文保留符号并用路径修补验证候选头。在六个主要Pythia(模型,任务)组合中,验证后的头分为两类:写手(直接促进正确标签)和消音者(效应相反)。路径修补提示下的标签预测可复现保留组的损伤结果,符号随机置换检验在五组中拒绝随机分配。未用于角色分配的测量也显示相同结构:写手更关注示范标签,消音者更关注格式标记,且其输出向量方向相对于同层控制项呈对抗偏移。按幅度排序的平均消融基线在层级任务中优先恢复消音者,在模块任务中恢复写手。所有十五个测试单元(六种Pythia规模、三种架构)均重复出现符号损伤方向。跨模板迁移表明这些角色为任务条件化,可维持、减弱或反转。零消融消音者后,六组主实验中正确标签对数似然差提升+0.13至+0.29纳特,准确率点估计全部提高。结果表明因果重要性与功能角色可分离,函数向量虽为有效任务表征,其头级实现却包含对抗性、任务依赖成分。

原文摘要 · Abstract (English)

Function-vector (FV) analyses commonly identify attention heads by the magnitude of their causal contribution to in-context tasks. Magnitude does not retain the direction of the effect on the task readout. We preserve the sign and validate candidate heads with path patching. Across six main Pythia (model, task) cells, the validated population separates into writers, whose direct effects favour the rule-correct label, and cancellers, whose effects oppose it. Labels assigned on path-patching prompts predict held-out group lesions, and a sign-shuffle null rejects a chance partition in five cells. Measurements not used to assign the roles show corresponding structure. Writers attend more to demonstration labels, cancellers more to format tokens, and their OV write directions are shifted toward opposition relative to same-layer controls. A magnitude-ranked mean-ablation baseline preferentially recovers cancellers on the hierarchical task and writers on the modular task. Signed lesion directions recur in all fifteen cells tested across six Pythia scales and three architectures. Cross-template transfer shows that these are task-conditioned roles that can persist, weaken, or reverse. Zero-ablating cancellers raises the correct-label logit difference by +0.13 to +0.29 nats in all six main cells, with accuracy point estimates increasing in all six. Together, the results separate causal importance from functional role and show that a function vector can remain a useful task-level representation while its head-level causal implementation contains opposed, task-conditioned components.

注意力头提示学习功能分解模型解释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。