arXiv:2603.00523cs.CLcs.AI2026-03被引 1

用多重剪枝评估电路稳定性,区分可靠结构与阈值噪声。

CIRCUS: Circuit Consensus under Uncertainty via Stability Ensembles

  • 通过多组剪枝配置计算边的出现频率,识别稳定连接。
  • 共识电路比所有配置并集小40倍,仍保持解释力。
  • 适合关注模型可解释性、避免人为阈值偏差的研究者。

每个机制电路都隐含一个不确定因素:它不仅反映模型计算,还受剪枝阈值选择的影响。改变阈值,电路即变,但现有方法将单一剪枝子图视为真实结构,无法区分稳健结构与阈值带来的伪影。我们提出CIRCUS,将电路发现重构为解释不确定性问题。CIRCUS在B种配置下对归因图进行剪枝,为每条边分配[0,1]范围内的经验包含频率s(e),衡量其跨配置的鲁棒性,并提取存在于所有视图中的共识电路。该方法实现原理化的核心/附带/噪声分解(类比贝叶斯变量选择中的后验包含指标),有效分离稳健结构与阈值敏感伪影,开销极低。在Gemma-2-2B和Llama-3.2-1B上,共识电路规模仅为所有配置并集的40倍,同时保留相当的影响力流动解释能力,持续优于影响度排序和随机基线,且通过激活补丁实验验证了其因果相关性。

原文摘要 · Abstract (English)

Every mechanistic circuit carries an invisible asterisk: it reflects not just the model's computation, but the analyst's choice of pruning threshold. Change that choice and the circuit changes, yet current practice treats a single pruned subgraph as ground truth with no way to distinguish robust structure from threshold artifacts. We introduce CIRCUS, which reframes circuit discovery as a problem of uncertainty over explanations. CIRCUS prunes one attribution graph under B configurations, assigns each edge an empirical inclusion frequency s(e) in [0,1] measuring how robustly it survives across the configuration family, and extracts a consensus circuit of edges present in every view. This yields a principled core/contingent/noise decomposition (analogous to posterior model-inclusion indicators in Bayesian variable selection) that separates robust structure from threshold-sensitive artifacts, with negligible overhead. On Gemma-2-2B and Llama-3.2-1B, consensus circuits are 40x smaller than the union of all configurations while retaining comparable influence-flow explanatory power, consistently outperform influence-ranked and random baselines, and are confirmed causally relevant by activation patching.

可解释性模型剪枝不确定性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。