arXiv:2606.16920cs.LGcs.AI2026-06被引 1

提出新方法降低大模型电路发现的不确定性,揭示其本质难点。

Demystifying Variance in Circuit Discovery of LLMs

论文配图:Demystifying Variance in Circuit Discovery of LLMs
图 1 · 摘自论文原文
  • 改进EAP-IG方法,理论保障下显著减少重采样偏差
  • 发现提示模板变化导致电路漂移,模型难被统一控制
  • 指出样本级不稳定性多由评估定义引发,非电路缺陷

电路发现是机制可解释性中定位模型关键组件的核心技术。尽管当前最先进的方法EAP-IG在(不)忠实性指标上表现良好,但仍存在显著变异性:包括重采样偏差(同一分布数据批次下电路变化)、改写偏差(提示重述时电路漂移)和样本级偏差(整体忠实性高但单样本波动大)。本文探究这些变异的根源。我们证明,新方法CEAP在理论保证下可大幅降低重采样偏差。进一步发现,不同提示模板会激活模型中不同电路,表明难以找到能覆盖所有表达形式的通用电路,暗示大模型可能本质上难以被精确引导。我们还发现,虽有研究声称稀疏性可生成更紧凑可读的电路,但无法解决该问题。对于样本级偏差,我们认为其大多无害:极低忠实性分数常源于忠实性定义方式,而非电路本身缺陷。我们指出,选择性贡献缩放这一神经机制会导致忠实性量级异常,解释了部分极端低分现象。

原文摘要 · Abstract (English)

Circuit discovery is a key technique in mechanistic interpretability to pinpoint the model components that are crucial for performing a given task. Although the current state-of-the-art method (EAP-IG) performs well on the metric of (un)faithfulness, it suffers from substantial variability. This includes resampling variance, where the circuit changes when we probe with a new batch of data from the same distribution; rephrasing variance, where the discovered circuit shifts when the prompts are rephrased; and sample-wise variance, where a circuit with low population unfaithfulness exhibits large fluctuations in unfaithfulness across individual samples. This paper studies the roots of these variances. We demonstrate that CEAP, our new circuit discovery method that improves upon EAP-IG with a theoretical guarantee, can substantially lessen resampling variance. We further show that rephrasing variance arises because prompts with different templates tend to activate different circuits in the model. This leads us to argue that it may be challenging to find a comprehensive circuit that explains and controls the model's behavior on a task, which can be expressed in countless templates, suggesting that LLMs may be inherently hard to steer. We show that sparsity, which has been claimed to form more compact and interpretable task circuits, fails to solve this problem. Regarding sample-wise variance, we argue that it is largely benign: extremely poor unfaithfulness scores often stem from how unfaithfulness is defined, rather than from defects in the measured circuits. We show that the magnitude of unfaithfulness is affected by selective contribution scaling, a neural mechanism that accounts for the extremely poor scores sometimes observed.

大模型可解释性电路发现变异性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。