发现大模型指令处理有局部非线性特征,能像开关一样选择不同路径。
Patches of Nonlinearity: Instruction Vectors in Large Language Models
- 通过因果中介分析,定位指令表示为局部向量(IVs)
- IVs具有线性可分但非线性交互的矛盾特性
- 提出新方法剥离线性假设,揭示其作为电路选择器的作用
尽管指令微调语言模型取得了显著成功并广泛应用,但其内部如何处理指令仍不清楚。本文从机制角度研究后训练阶段(监督微调SFT与直接偏好优化DPO)中指令表征的构建与利用。通过因果中介分析,发现指令表征在模型中较为局部。这些表征被称为指令向量(IVs),表现出线性可分与非线性因果交互的奇特共存,对机械可解释性中常见的线性表征假说提出广泛质疑。为解耦非线性因果交互,提出一种不依赖补丁技术隐含线性假设的新方法。结果表明,在早期层形成任务表征后,后期层会根据该表征选择不同信息路径来完成任务,即IVs充当电路选择器。
原文摘要 · Abstract (English)
Despite the recent success of instruction-tuned language models and their ubiquitous usage, very little is known of how models process instructions internally. In this work, we address this gap from a mechanistic point of view by investigating how instruction-specific representations are constructed and utilized in different stages of post-training: Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO). Via causal mediation, we identify that instruction representation is fairly localized in models. These representations, which we call Instruction Vectors (IVs), demonstrate a curious juxtaposition of linear separability along with non-linear causal interaction, broadly questioning the scope of the linear representation hypothesis commonplace in mechanistic interpretability. To disentangle the non-linear causal interaction, we propose a novel method to localize information processing in language models that is free from the implicit linear assumptions of patching-based techniques. We find that, conditioned on the task representations formed in the early layers, different information pathways are selected in the later layers to solve that task, i.e., IVs act as circuit selectors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。