arXiv:2410.01434cs.LGcs.CL2024-10ACL被引 7

发现语言模型中可复用的计算模块,能组合实现复杂任务。

Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models

  • 基于上下文无关语法,识别十种字符串编辑操作的计算电路。
  • 功能相似的电路存在显著节点重叠与跨任务一致性。
  • 电路可通过集合运算组合,构建更复杂的模型能力。

解释性研究的核心问题之一是神经网络(尤其是语言模型)是否通过子网络实现可复用的功能,并能组合完成更复杂的任务。近期机制可解释性进展已识别出‘电路’——即负责模型在特定任务上行为的最小计算子图。然而,多数研究集中于单一任务的电路识别,缺乏对功能相似电路之间关系的探讨。为填补这一空白,我们通过分析基于Transformer的语言模型中高度组合性子任务的电路,研究其模块化特性。具体而言,基于概率上下文无关语法,我们识别并比较了负责十种模块化字符串编辑操作的电路。结果表明,功能相似的电路表现出显著的节点重叠和跨任务忠实性。此外,我们证明所识别的电路可通过集合运算进行复用与组合,从而表征更复杂的模型功能。

原文摘要 · Abstract (English)

A fundamental question in interpretability research is to what extent neural networks, particularly language models, implement reusable functions through subnetworks that can be composed to perform more complex tasks. Recent advances in mechanistic interpretability have made progress in identifying $\textit{circuits}$, which represent the minimal computational subgraphs responsible for a model's behavior on specific tasks. However, most studies focus on identifying circuits for individual tasks without investigating how functionally similar circuits $\textit{relate}$ to each other. To address this gap, we study the modularity of neural networks by analyzing circuits for highly compositional subtasks within a transformer-based language model. Specifically, given a probabilistic context-free grammar, we identify and compare circuits responsible for ten modular string-edit operations. Our results indicate that functionally similar circuits exhibit both notable node overlap and cross-task faithfulness. Moreover, we demonstrate that the circuits identified can be reused and combined through set operations to represent more complex functional model capabilities.

可解释性模块化语言模型电路

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。