arXiv:2606.08292cs.AI2026-06

测试注意力头能否跨提示迁移计算,发现多数无法实现干净转移。

Necessary, Decodable and Reversible, Yet Not Transferable: A Stress Test for Attention-Head Role Claims

  • 通过匹配控制实验检验注意力头状态是否可跨提示迁移计算。
  • 在15个模型组件中仅8个表现出计算转移,且多为惰性或泛化效应。
  • 结果表明必要性、可解码性与跨提示迁移是独立证据,不互为充分条件。

机制研究常通过移除组件导致行为受损、激活线性编码任务信息、恢复激活可修复损伤来赋予组件特定角色。本文测试更强假设:同一激活能否将计算传递至另一提示?在覆盖三个7-8B指令微调模型和五个单步计算家族的15个模型组件中,仅有约8个通过了匹配控制转移测试,未实现清晰计算迁移。相同答案与匹配上下文对照组揭示了惰性或广泛影响。正向对照验证了工具在未指定提示中可恢复已知功能向量,并在三个受控模型种子之一中恢复紧凑可覆盖的残差状态,但均不能证明对真实模型注意力在提示覆盖时所承载信息的敏感性。在Qwen的两种组合形式中,虽存在强可解码操作方向,但在通过门控条件下仍无法引导生成答案;跨模型测试无效。结果表明必要性、可解码性、同提示修复与跨提示迁移是可分离的证据,不共同构成确凿结论。不能排除未测试组件或场景中存在可移植选择状态的可能性。

原文摘要 · Abstract (English)

Mechanistic studies often assign a component a role when removing it damages a behavior, its activation linearly encodes task information, and restoring that activation repairs the damage. We test the stronger implication: can the same activation carry the requested computation into another prompt? In the roughly 8 of 15 model-family cells receiving our full matched-control transfer assay, covering three 7-8B instruction-tuned models and 5 single-step computation families, tested attention-head states did not produce clean computation transfer. Same-answer and matched-context controls instead exposed inert or broad effects. Positive controls bound this result rather than eliminating its caveats: the instrument recovers known function vectors when writing into an underspecified prompt, and recovers a compact override-capable residual state in one of three controlled-model seeds. Neither control proves sensitivity to the object carried by real-model attention during prompt override. On two composed forms in Qwen, a strongly decodable operation direction also failed to steer the generated answer under passing gates; the corresponding cross-model test was invalid. These results show that necessity, decodability, same-prompt repair, and cross-prompt transfer are separable evidence. They do not establish that a portable selection state is absent from untested components or settings.

注意力机制可解释性模型测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。