用归纳逻辑构建可累积的神经网络机制理论,让电路发现可比较、可迁移。
From Circuit Evidence to Mechanistic Theory: An Inductive Logic Approach

- 将电路行为分解为因果功能签名和架构签名,形成统一形式化表示
- 揭示不同任务类型采用不同计算策略,如注意力复制与MLP绑定
- 支持跨模型规模和架构家族的原理性迁移,优于传统图核方法
机制可解释性虽能实现神经网络行为的电路级因果分析,但发现的电路常呈孤立实验现象:缺乏共享的形式化表达来描述电路计算内容、相互关系,或判断两个发现是否支持同一机制。本文通过将电路解释视为归纳理论构建,提供累积性机制科学的形式基础设施。每个电路在两个层面表征:因果功能签名(CFS),基于因果归因证据和标记角色分布;以及由归纳逻辑编程(ILP)从尺度不变结构谓词中学习的架构签名τₐᵣcₕ。二者构成形式化一致性层,使机制主张显式、可通过θ-包含比较,并跨模型规模可移植。CFS揭示任务类型间存在质异的计算策略,如注意力介导的复制与MLP介导的绑定。ILP签名在结构分离上显著优于图核和特征向量基线,支持原则性跨规模与架构家族迁移。
原文摘要 · Abstract (English)
Mechanistic interpretability produces circuit-level causal analyses of neural network behaviour, but discovered circuits often remain isolated experimental artefacts: there is no shared formal representation for what circuits compute, how they relate, or when two findings provide evidence for the same mechanism. This work provides a formal infrastructure for cumulative mechanistic science by treating circuit interpretation as inductive theory construction. Each circuit is characterised at two levels: a Causal Functional Signature (CFS), which grounds component behaviour in causal attribution evidence and token role profiles, and an architectural signature $τ_{\mathrm{arch}}$, learned by inductive logic programming (ILP) from scale-invariant structural predicates. Together, these constitute a formal coherence layer that makes mechanistic claims explicit, comparable via $θ$-subsumption, and portable across model scales. CFS reveals qualitatively distinct computational strategies across task types, including attention-mediated copying versus MLP-mediated binding. ILP signatures achieve substantially better structural separation than graph kernel and feature-vector baselines, and support principled transfer across model scales and architecture families.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。