测量模型合并干扰的指标其实看的是门控和预算,而非真实干扰。
Reading the Gate, Not the Interference: Output-Side Interference Measurement Does Not Track Merge Collapse
- 用精确层间激活交叉项分析合并后的干扰机制。
- 交叉项在合并崩溃时是旁观者,不影响实际性能下降。
- 适合研究模型合并失效机理的研究者阅读。
任务算术合并在特定条件下会失效,现有方法通过测量合并模型内部干扰来诊断原因。本文采用最直接的度量方式——因子分解账本中的精确层间激活交叉项,揭示其因果结构:每个模块主要传递并放大交叉项,而非生成它;若被移除,仅剩99%的范数可由未受损的边际路径恢复,除非后期移除;其输出效应随位移角度单调变化(正交位移加剧干扰);基于两假设的模型可推导出角度规律,并回溯预测剂量曲线(R² ≥ 0.99)。该度量所追踪的并非领域假设的内容:行为专家性与该度量解耦,跨条件表现受分母主导——指令模板将主效应控制在1%以内,而绝对交互增长111倍(从2个到6个任务),导致在k=2时抑制表达干扰,在k=6时放大干扰。在合并真正崩溃时,交叉项为旁观者,即使全程擦除也无法修复崩溃。此时输出侧比值不携带方法信息,而已有两种状态空间度量在两个尺度上均正确排序。所有81个预测在数据前已冻结,被证伪的结果如实报告。输出侧干扰测量读取的是门控、分母与位移预算,而非真正的干扰。失败的根源在于载体与旁观者的混淆:崩溃源于边际位移,交叉项仅伴随其发生,唯有状态空间能识别载体。
原文摘要 · Abstract (English)
Task-arithmetic merging works until it doesn't, and the field diagnoses why by measuring interference inside the merged model. We take the most direct such measure, the exact layerwise activation cross-term of a factorial ledger, establish its causal anatomy, and then ask what it tracks. The anatomy is clean: each block mostly transports and amplifies the cross-term rather than generating it; erased, it is regenerated by the untouched marginal paths to 99% of its norm unless removed late; its output effect varies monotonically with the displacement's angle (orthogonal displacements make interference worse), and a two-assumption model derives the angle law and retro-dicts the dose curve (R^2 >= 0.99). What the measure tracks is not what the field assumes. Behavioural expert-likeness is decoupled from it across four instruments. Its cross-condition behaviour is denominator-dominated: an instruction template pins the main effect to within 1% while the absolute interaction grows 111x from two to six merged tasks, suppressing expressed interference at k=2 and amplifying it at k=6. And where merging actually collapses, the cross-term is a bystander, not the carrier: across two collapse parameterizations at two scales, even erased persistently at every position, removing it entirely repairs none of the collapse. There the output-side ratio carries no method information under a common counterfactual, while two state-space measures the field already uses rank methods correctly at both scales. All 81 predictions were frozen before their data; falsifications are reported as such. Output-side interference measurement reads the gate, the denominator, and the displacement budget, not the interference. What fails a merge is the carrier-bystander split: collapse rides in the marginal displacements while the cross-term merely accompanies it, and only state space sees the carrier.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。