arXiv:2608.03629cs.AIcs.LG2026-08被引 1

揭示跨层交互的数学边界,验证大模型真实权重下的注意力机制行为

Cross-Layer Interaction under Weight-Space Ablation: A Closed-Form Attention Jacobian Bound and a Test on a Real Pretrained Model

  • 提出跨层扰动的精确分解公式,分离同层与跨层交互项
  • 推导注意力子模块的雅可比界,实测未在Qwen2.5-1.5B中出现违反
  • 发现隐含间接宾语识别电路,跨层交互在三组实例中可观测

本文扩展了先前研究在理想化模型中关于激活修补与权重扰动一致性的问题。针对多层架构,首次将跨层交互精确分解为各残差块内项与一个不保证小量的跨层余项。进一步,对两层情形精确表征该余项为混合二阶导数的双重积分,并提出需补全的注意力子块雅可比界。本研究以闭式推导出该界,并在真实Qwen2.5-1.5B-Instruct权重上验证无一例外。同时,闭式给出伴随论文未揭示的曲率常数。最后,在同一模型上使用原始激活修补法探测到一个未预设的间接宾语识别电路;测试显示五例中共享载体存在,多数案例满足坍缩与解离,且三组跨块层对间可测量非零交互。

原文摘要 · Abstract (English)

A companion paper studies when activation patching and weight-space ablation agree, inside an idealized model where a conditional computation is carried additively through a residual stream. For the one composition in that model where two carriers are architecturally dependent, an attention head and its own layer's normalization-MLP composition, it derives an exact first-order interaction formula, zero when only the MLP is ablated and second-order bounded when the head is also ablated. That result is confined to a single residual block and checked only on small transformers on a synthetic task. This paper extends the result past both limits. First, the interaction from ablating carriers spanning several layers decomposes exactly into same-block terms, one per touched layer, plus a cross-layer remainder on which the decomposition makes no claim of smallness. Second, we isolate that remainder exactly, for two layers, as a double integral of a mixed second derivative, and name the missing ingredient needed to bound it: a Jacobian bound for the attention sub-block. We derive this bound in closed form and verify it, without a single violation, against Qwen2.5-1.5B-Instruct's real weights, though we do not yet chain it across layers. We also give, in closed form, the curvature constant the companion paper's bound leaves unexhibited. Third, on that same model, we search for and find an emergent circuit for indirect object identification, never designed into it, using the original activation-patching method for this task, and test collapse, dissociation, and interaction on it. The result is mixed: a shared carrier emerges across all five tested instances, collapse and dissociation hold on most but not all, and a nonzero interaction is measurable on three of five, at layer pairs outside the same-block case the companion theorem covers.

注意力机制模型可解释性雅可比分析大模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。