不同模型用相同方法找因果回路,结果却各不相同,说明模式选择性不等于任务因果结构。
Pattern Selectivity is Not Task-Causal Structure: A Cross-Architecture Mechanistic Study of Composed-Task Circuits in 1B-Class Language Models
- 用统一流程测试三类10亿级模型的注意力回路,发现同任务对应不同模式
- 12个实验单元中无两个共享相同主要因果信号,验证了结果的模型特异性
- 提出五类因果归因分类,支持未来可验证的混合专家模型机制假设
我们检验单一‘筛选-消融’方法(通过任务模式选择性识别注意力头回路,再用匹配随机空模型验证因果性)在不同模型族间是否产生一致的机械解释。该方法虽可跨训练管道移植,但具体识别出的回路并不一致。在四个组合任务(间接宾语识别、大于比较、后继序列、变量绑定)与三种10亿级语言模型(Pythia 1B / Pile / dense;OLMo 1B / DCLM / dense;OLMoE 1B-7B / DCLM / 混合专家)上,采用统一协议,匹配随机空模型在每组重复十次。12个(任务, 模型)组合中,没有两组在可比效应量下拥有相同的主因筛选结果:相同任务、相同行为能力,不同模型使用不同注意力模式实现。我们引入五类筛选-结果分类(主因、次因、相关、干扰、空)并设定量化阈值,结果显示五类均出现。提出可证伪假说:本研究中的混合专家模型在前一词位置基底结构上构建组合任务回路(对OLMoE 1B-7B的3个任务,前词回路消融为最强因果筛选),而间接宾语识别例外符合其为末位名称复制任务、直接探测不同模式的特性。该假说对其他混合专家模型有明确预测。
原文摘要 · Abstract (English)
We test whether a single screen-and-ablate recipe -- identify attention-head circuits by task-pattern selectivity, then verify by causal ablation against a matched-random null -- produces consistent mechanistic claims across model families. The recipe ports across pipelines; the specific circuit it identifies does not. Across four composed tasks (indirect-object identification, greater-than, successor sequences, variable binding) and three 1B-class language models from distinct training pipelines (Pythia 1B / Pile / dense; OLMo 1B / DCLM / dense; OLMoE 1B-7B / DCLM / mixture-of-experts), we run a unified protocol with the matched-random null sampled across ten seeds per cell. The resulting 12 (task, model) cells contain no two that share the same primary causal screen at comparable effect size: the same task, with the same behavioral capability, is implemented through different attention-pattern types across models. We introduce a five-category screen-outcome taxonomy -- primary cause, secondary cause, correlate, interferer, null -- with quantitative thresholds, and show that all five outcomes appear in the panel. We propose a falsifiable hypothesis: the MoE model in our panel builds composed-task circuits on top of a foundational previous-token positional substrate (the prev-token-circuit ablation is the strongest causal screen on 3 of 4 tasks for OLMoE 1B-7B), with the IOI exception consistent with IOI being a final-position name-copying task whose structure directly probes a different pattern. The hypothesis comes with explicit predictions for other MoE language models. We frame the methodology honestly: the spectral participation-ratio signal from the companion methodology paper is a general indicator of specialized computation; what makes a finding task-specific is the task-pattern screen plus a per-model causal verification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。