通过稀疏神经元干预,精准激活语言模型的任务行为。
Distributed Sparse Interventions in Language Models

- 在神经元层面进行稀疏干预,捕捉非线性效应。
- 仅需0.01%神经元即可激活任务行为。
- 适合研究模型内部机制与任务分解的学者。
语言模型能以不同抽象层次灵活处理任务,具备从上下文推断、并行执行及选择任务的能力。为探究模型组件对任务行为的影响,可通过干预分析其因果作用。以往模型调控研究多集中于激活空间中的全局方向,将任务表示近似为线性和可加。通过神经元级别的干预研究,我们发现显著且神经元特异的非线性效应,现有方法无法捕捉。为此提出分布式稀疏干预(DSI),考虑跨层神经元间的非线性与交互关系,识别出稀疏的神经元集合以引发任务相关计算。在多种任务中,DSI仅需干预0.01%的神经元即可激活指令微调语言模型中的任务行为,证明稀疏分布式干预在神经元基下的有效性。此外,采用集合视角使对识别出的神经元集进行计算操作成为可能,通过分析其在多任务中的影响,揭示单个神经元的作用。该方法实现对模型行为的细粒度控制、任务相关神经元集的定位,并深化对任务构成的理解。
原文摘要 · Abstract (English)
Language models perform a wide range of tasks at varying levels of abstraction with the capacity to flexibly infer tasks from context, execute multiple tasks simultaneously, and select among competing tasks. To study the role of model components in task behaviour, their causal influence can be investigated through interventions. Prior work on model steering has largely focused on interventions along global directions in activation space, modeling task representations as approximately linear and additive. By studying interventions at the neuron level, we find substantial, neuron-specific nonlinear effects on model outputs that are not captured by current steering approaches. We introduce Distributed Sparse Interventions (DSI), an intervention approach that considers nonlinearities and interactions between neurons across layers to identify sparse sets of neurons that elicit task-relevant computations. Across a range of tasks, we demonstrate that DSI can activate task behaviour in instruction-tuned language models by localising and intervening on as few as 0.01% of neurons, highlighting the effectiveness of sparse, distributed interventions in the neuron basis. Additionally, adopting a set-based perspective enables computations over the identified neuron sets, offering insights into the roles of individual neurons by analysing their effects across tasks. Through sparse interventions, DSI enables fine-grained control over model behaviour, localisation of task-relevant neuron sets, and furthers our understanding of task composition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。