arXiv:2606.20678stat.MLcs.LG2026-06被引 3

新方法区分注意力机制中内容传递与计算破坏,发现早期通道负责关系信息传输。

Beyond Importance: Interchange-Sobol Sensitivity Reveals Task-Specific Content Channels in Transformer Components

  • 用替换与删除对比的干预框架分离内容传输与计算依赖
  • 在GPT-2和Qwen2.5-1.5B中发现早期层内容通道被传统方法低估
  • 揭示了第9层第8头专门处理关系信息,适合模型可解释性研究者

现有机制可解释性方法通过单一重要性得分总结Transformer组件,混淆了两种不同作用:组件的重要性可能源于其传递任务相关内容的能力,或因移除其贡献导致前向计算退化。本文提出互换-组Sobol分解(IGSD),一种成对干预框架,通过比较相同组件的激活替换与零消融,估计两个类似Sobol的方差指数,并用其带符号差值分离这两种角色,同时通过对称离流形诊断 $\ ext{ST}>1$ 监控干预有效性。在事实回忆任务中,IGSD识别出GPT-2 small与Qwen2.5-1.5B中早期层的内容通道,而标准重要性方法低估了该通道。受控的主体与关系供体设计表明,早期通道传输关系-框架内容,而后期注意力负责主体检索内容,细化至已知的 $\\mathrm{Attn}_{L9H8}$ 头。后期层钳制实验确认,早期信号通过下游变换表达,而非残差直通。结果表明,替换与删除并非等价控制,其差异为内容传输提供了实用的统计诊断。

原文摘要 · Abstract (English)

Mechanistic interpretability methods summarize a transformer component by a single importance score, conflating two distinct roles: a component may matter because it transports task-relevant content, or because the forward computation degrades when its contribution is removed. We introduce \emph{Interchange-Group Sobol Decomposition} (IGSD), a paired-intervention framework that compares matched activation replacement with zero ablation on the same component, estimates two Sobol-style variance indices, and uses their signed difference to separate the two roles, with intervention validity monitored by a symmetric off-manifold diagnostic $\widehat{\mathrm{ST}}>1$. In factual recall, IGSD identifies an early-layer content channel in both GPT-2 small and Qwen2.5-1.5B that standard importance methods underestimate. A controlled subject and relation donor design shows that the early channel transports relation-frame content while late attention transports subject-retrieval content, refining at head granularity to the known $\mathrm{Attn}_{L9H8}$ head. Late-layer clamping confirms that the early signal is expressed through downstream transformations rather than residual pass-through. These results show that replacement and deletion are not interchangeable controls and their divergence provides a practical statistical diagnostic for content transport in transformer components.

可解释性Transformer注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。