arXiv:2411.16105cs.LGcs.AI2024-11被引 8

研究大模型电路如何适应不同提示格式,发现其核心机制高度通用。

Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability

  • 通过测试提示变体,分析IOI电路的组件复用与行为变化。
  • 电路在多种挑战性提示中仍有效,仅新增输入边,未改变核心机制。
  • 发现新机制S2 Hacking,解释为何原算法失效时仍能工作,适合模型可解释性研究者。

机制可解释性旨在通过识别神经网络中的电路(即执行特定任务的最小子图)来理解大型神经网络的内部运作。这些电路通常在特定提示格式下被发现和分析。然而,由于大语言模型(LLMs)能在不同提示格式间泛化同一任务,现有电路的泛化能力尚不明确。例如,模型泛化是复用相同电路组件、组件行为改变,还是使用全新组件?本文研究了GPT-2 small中已深入研究的间接对象识别(IOI)电路的泛化能力,该电路被认为实现了一个简单可解释的算法。我们在挑战该算法假设的提示变体上评估其性能。结果表明,该电路表现出意外的强泛化能力:完全复用原有组件与机制,仅新增输入边。值得注意的是,即使在原始算法应失效的提示变体中,电路仍有效;我们发现一种名为S2 Hacking的新机制解释此现象。结果表明,大模型中的电路可能比先前认知更灵活且更具泛化性,强调了研究电路泛化对理解模型整体能力的重要性。

原文摘要 · Abstract (English)

Mechanistic interpretability aims to understand the inner workings of large neural networks by identifying circuits, or minimal subgraphs within the model that implement algorithms responsible for performing specific tasks. These circuits are typically discovered and analyzed using a narrowly defined prompt format. However, given the abilities of large language models (LLMs) to generalize across various prompt formats for the same task, it remains unclear how well these circuits generalize. For instance, it is unclear whether the models generalization results from reusing the same circuit components, the components behaving differently, or the use of entirely different components. In this paper, we investigate the generality of the indirect object identification (IOI) circuit in GPT-2 small, which is well-studied and believed to implement a simple, interpretable algorithm. We evaluate its performance on prompt variants that challenge the assumptions of this algorithm. Our findings reveal that the circuit generalizes surprisingly well, reusing all of its components and mechanisms while only adding additional input edges. Notably, the circuit generalizes even to prompt variants where the original algorithm should fail; we discover a mechanism that explains this which we term S2 Hacking. Our findings indicate that circuits within LLMs may be more flexible and general than previously recognized, underscoring the importance of studying circuit generalization to better understand the broader capabilities of these models.

机制可解释性大模型泛化电路分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。