发现大模型任务可由多种不同结构电路实现,打破单一机制假设。
All Circuits Lead to Rome: Rethinking Functional Anisotropy in Circuit and Sheaf Discovery for LLMs

- 通过结构去重方法发现多个无重叠的高效功能电路
- 验证单个任务存在多个等效且不依赖特定边的稀疏电路
- 适合关注模型可解释性与机制多样性研究者阅读
本文通过实证与理论证据质疑了大语言模型电路与层析发现(CSD)中的核心隐含假设——功能各向异性假说:即模型功能由唯一或近似唯一的内部机制承载。我们发现,同一任务可由多个结构各异但均忠实、稀疏且完整的电路或层析同时支持。为此,提出「重叠感知层析排斥」方法,在目标中显式惩罚多轮发现中的结构重叠,从而揭示出在多种常见基准上表现优异但结构几乎不共享的机制。随着发现的层析数量增加,此现象愈发显著,并在主流CSD方法中保持稳健。我们进一步识别出一个超稀疏的三边层析,其任意单边均非必需,动摇了对标准或关键组件的弱化认知。为此提出分布式密集电路假说,并通过理论分析表明:在高维叠加下,非唯一且低重叠的电路解释可在温和假设下自然产生。结果表明,大模型的机制解释本质上是非标准的,需重新审视CSD结果的解读与评估方式。
原文摘要 · Abstract (English)
In this paper, we present empirical and theoretical evidence against a central but largely implicit assumption in circuit and sheaf discovery (CSD), which we term the Functional Anisotropy Hypothesis: the idea that functions in large language models (LLMs) are localised to a unique or near-unique internal mechanism. We show that a single LLM task can instead be supported by multiple, structurally distinct circuits or sheaves that are simultaneously faithful, sparse, and complete. To systematically uncover such competing mechanisms, we introduce Overlap-Aware Sheaf Repulsion, a method that augments the CSD objective with an explicit penalty on structural overlap across multiple discovery runs, enabling the discovery of circuits or sheaves with strong task performance but minimal shared structure across a plethora of common CSD benchmarks. We find that this phenomenon becomes increasingly pronounced as the number of discovered sheaves grows and persists robustly across major CSD methods. We further identify an ultra-sparse three-edge sheaf and show that none of its edges is individually indispensable, undermining even weakened notions of canonical or essential components. To explain these findings, we propose a Distributive Dense Circuit Hypothesis and provide a theoretical analysis demonstrating that non-unique, low-overlap circuit explanations arise naturally from high-dimensional superposition under mild assumptions. Together, our results suggest that mechanistic explanations in LLMs are inherently non-canonical and call for a rethinking of how CSD results should be interpreted and evaluated.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。