arXiv:2608.11767cs.CL2026-08

语言模型用类型分组路由因果知识,但不直接决定答案输出。

Causal Structure is Inducible but Functionally Decoupled: The Routing/Readout Boundary of a Typed Mechanism Library

  • 按证据类型分槽的机制库通过监督信号自动形成路由结构。
  • 路由与答案读出严格分离,误差低于3.4×10⁻⁶,无副作用。
  • 结构零成本且可精确回滚,适合可解释性研究者使用。

当语言模型回答干预性问题时,其计算依赖于查询所需的证据类型。我们发现变压器在组织因果知识时存在一种解耦:由类型级监督诱导的按类型分槽结构负责路由,却与答案读出功能完全解耦。在包含确切干预真值的因果世界基准上,采用冻结协议,在2260万和1.25亿参数规模下验证了这一现象。四项预注册发现:(i) 源头。按类型分槽结构由类型级监督诱导,在结构相同但无监督的对照组中不存在,也无法通过内容无关的门控标签获得,统计上可归因于监督信号,1.25亿规模下预注册协议验证通过九个单元。(ii) 边界。该结构是清晰的类型路由索引,具有锐利的路由/读出边界:槽码仅支撑路由,不驱动答案读出(|Δŷ| ≤ 3.4×10⁻⁶,零副作用,三种子,跨5.6倍规模窗口稳定),因此不主张行为可编辑性。(iii) 成本。该结构零成本:模型质量与参数匹配的单体模型相差不超过0.0082纳特。(iv) 信任。机制库状态在编辑下局部精确,比特级可回滚——每种子执行250次单次编辑与1000次堆叠回滚,零失败。此外发现,无监督基线本身随规模变化,因此在不同规模间复用校准过的基线可能产生混淆。所有结论均基于预先注册、机器可验证的标准,审计轨迹(含一项未通过标准及冻结协议处理方式)已作为附录公开。

原文摘要 · Abstract (English)

When a language model answers an interventional question, the computation it must perform depends on the type of evidence the query requires. We report a decoupling in how a transformer organizes causal knowledge: slot-by-type structure induced by type-level supervision organizes routing, yet remains functionally decoupled from answer readout. We establish this with a typed mechanism library -- discrete mechanism slots partitioned by evidence type, auditable at the state level -- on a causal-world benchmark with exact interventional ground truth, under a frozen protocol, at two scales (22.6M and 125M). Four preregistered findings. (i) Origin. Slot-by-type organization is induced by type-level supervision: absent in architecturally identical unsupervised controls, not buyable by content-free gating labels, and statistically attributable to the supervision signal, replicating at 125M under a powered preregistered protocol (all nine cells passed). (ii) Boundary. The induced structure is a typed routing index with a sharp routing/readout boundary: slot codes scaffold routing but do not drive answer readout ($|Δ\hat{y}| \le 3.4\times10^{-6}$, zero collateral, three seeds, stable across a 5.6x scale window) -- we therefore make no behavioral-editability claim. (iii) Cost. The structure is free: LM quality matches a parameter-matched monolith within 0.0082 nats. (iv) Trust. The library state is exactly local under edit and bit-exactly revertible -- 250 single-edit and 1,000 stacked reverts per seed, zero failures. We further find that the unsupervised null itself moves with scale, so comparisons reusing a null calibrated at one scale may be confounded at another. Every claim is tied to a preregistered, machine-checkable criterion archived before the data it governs; the full audit trail, including one criterion we failed and how the frozen protocol handled it, is released as an appendix.

因果推理可解释性模型机制路由结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。