arXiv:2608.03620cs.LGcs.AI2026-08被引 1

揭示了模型权重删减与激活修补在条件计算中的分歧机制

A Theory of Conditional Collapse under Low-Rank Weight-Space Ablations: I. The Single-Block Theory and Synthetic Validation

  • 通过理想化残差流模型,推导出条件坍缩的精确数学条件
  • 实验证明删减与修补对输出的影响方向相反,且互不包含
  • 在合成任务中验证理论,相关性高达Spearman -0.83,适合因果分析研究者

激活修补和权重空间删减均声称某组件对行为有因果作用,但分别作用于前向传播结果与所有前向传播背后的参数。我们探讨二者何时一致。研究一个理想化模型:条件计算通过残差流线性叠加,$F(x)=F_0(x)+\sum_iα_i(x)v_i$,由线性泛函读出。证明三个精确结论:第一,删除一组载体仅当对输入对对称且无外部对比时,才会使匹配输入对坍缩至相同无条件输出,误差为确定性且可精确表达;第二,修补载体移动读出量为其源-目标对比值,而删减则移动其绝对水平,二者互不约束,可构造单个修补翻转决策但单个删减无法实现的情况;第三,对注意力头与其层归一化及MLP组合,推导出带可证二阶余项的一阶交互公式,仅当仅删去MLP时余项恒为零,一般情况下并非如此。在合成条件任务上训练的小型Transformer验证全部三类预测:39种删减配置下测量交互与理想模型预测准确率强秩相关(Spearman $-0.83$),另一任务与架构复现相同模式,包括极性反转。单块交互结果可推广至多残差块,合成验证亦扩展至真实预训练模型,详见配套论文。

原文摘要 · Abstract (English)

Activation patching and weight-space ablation both claim a component is causally responsible for a behavior, yet they act on different objects: one forward pass versus the parameters behind every forward pass. We ask when they agree. We study an idealized model where a conditional computation is carried additively through a residual stream, $F(x)=F_0(x)+\sum_iα_i(x)v_i$, read out by a linear functional, and prove three exact results. First, deleting a subset of carriers collapses a matched input pair onto the same unconditional output \emph{if and only if} the removal is symmetric on the pair and leaves no outside contrast; the error is deterministic, and we give its exact form even when the two conditions hold only approximately. Second, patching a carrier moves the readout by its donor-receiver \emph{contrast}, while ablating it moves the readout by its \emph{absolute level}; neither bounds the other, and we construct pairs where every single-carrier patch flips the decision while no single-carrier ablation does. Third, for an attention head composed with its own layer's normalization and MLP, we derive an exact first-order interaction formula with a provably second-order remainder, vanishing identically when only the MLP is ablated but not, in general, when a head is. Small transformers trained on a synthetic conditional task illustrate all three predictions: across thirty-nine ablation configurations the measured interaction is strongly rank-correlated with the idealized model's predictive accuracy (Spearman $-0.83$), and a second task and architecture reproduces the same pattern, including a further polarity reversal. The single-block interaction result extends past one residual block, and the synthetic validation is tested against a real pretrained model, in a companion paper that takes this theory further along both axes.

模型因果权重删减激活修补理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。