发现结构不同的电路可能执行相同功能,揭示了电路发现中的‘虚假特化’现象。
Many Circuits, One Mechanism: Input Variation and Evaluation Granularity in Circuit Discovery

- 通过改变输入词频测试电路发现,结构差异不等于机制不同。
- 75个电路中核心路径在不同频率下均能恢复99%以上性能。
- 需用边缘级评估和跨条件迁移验证才能避免误判机制差异。
电路发现方法旨在识别解释模型行为的子图,通常将发现的结构差异视为不同机制的证据。本文在固定任务下改变输入词频,发现电路结构上呈现频率特异性,但功能与表征分析未显示可靠差异,称为‘虚假特化’。在四个频率区间及加权控制条件下,从五个Pythia模型(70M-1.4B)中提取75个电路。结果显示,结构差异的电路实现相同计算:特定频段边在各频段间广泛转移,核心共享路径在70M以上模型中恢复至少99%性能,因果交换实验确认内部表示可互换。小规模主谓一致任务复现相同模式:结构差异但跨频段可迁移,多数频段共享核心恢复近全部准确率。同一频段内重复提取表明,发现算法在有效子图等价类中采样而非唯一机制。标准评估方式掩盖此现象:源级评估夸大忠实性,而边级评估揭示结构到功能的多对一映射。未测试组件内词位或特征层面的特化。结果表明,电路结构差异不足以证明机制不同,需边级评估与跨条件迁移测试来暴露真实关系。
原文摘要 · Abstract (English)
Circuit discovery methods identify subgraphs that explain model behaviors, and structural differences between discovered circuits are commonly interpreted as evidence of distinct mechanisms. We test this assumption by varying input-token frequency while holding the task fixed. The discovered circuits appear specialized by frequency when compared structurally, but functional and representational analyses show no reliable evidence of corresponding differences. We term this mismatch phantom specialization. Using the Literal Sequence Copying task across four frequency bands plus a frequency-weighted control, we extract 75 circuits from five Pythia models (70M-1.4B). We find that structurally distinct circuits implement the same computation: band-specific edges transfer broadly across bands, a core shared across most bands recovers at least 99% of circuit performance in models above 70M, and causal interchange interventions confirm that internal representations are interchangeable across frequency bands. A smaller subject-verb agreement replication shows the same pattern: circuits differ structurally but transfer broadly across bands, and the core shared by most bands recovers nearly all circuit accuracy. Repeated extractions within the same band further suggest that discovery algorithms sample from an equivalence class of valid subgraphs rather than recovering a unique mechanism. Standard evaluation practice obscures this pattern: source-level evaluation inflates apparent faithfulness, while edge-level evaluation reveals the many-to-one mapping from structure to function. We do not test specialization at the level of token positions or features inside a component. Our results show that structural differences between circuits are not sufficient evidence for distinct mechanisms, and that exposing this requires edge-level evaluation and cross-condition transfer tests.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。