通过线性路径提取语言模型中的独立技能,更精准揭示其内在机制。
Skill Path: Unveiling Language Skills from Circuit Graphs
- 将电路图分解为线性链式结构,隔离出单一技能路径。
- 实验证明三种通用语言技能具有分层与包容特性。
- 适合研究模型可解释性、内部机制的学者使用。
电路图发现已成为揭示语言模型技能机制的基础方法。尽管电路图能保持输出忠实性,但其原子消融问题会导致连接组件间因果依赖的丢失。此外,为保留输出忠实性而设计的发现过程,意外捕捉了除目标技能外的额外影响。为此,我们提出技能路径,通过在组件线性链中隔离个体技能,提供更精细、紧凑的表示。为从电路图中提取技能路径,我们提出三步框架:分解、剪枝与后剪枝因果中介分析。特别地,我们实现了对Transformer模型的完整线性分解,生成解耦计算图。剪枝后,进一步采用反事实与干预等因果分析技术,从电路图中提取最终的技能路径。为凸显技能路径的重要性,我们使用该框架研究了三种通用语言技能——前项标记技能、归纳技能和上下文学习技能。实验支持这些技能的两个关键性质:分层性与包容性。
原文摘要 · Abstract (English)
Circuit graph discovery has emerged as a fundamental approach to elucidating the skill mechanistic of language models. Despite the output faithfulness of circuit graphs, they suffer from atomic ablation, which causes the loss of causal dependencies between connected components. In addition, their discovery process, designed to preserve output faithfulness, inadvertently captures extraneous effects other than an isolated target skill. To alleviate these challenges, we introduce skill paths, which offers a more refined and compact representation by isolating individual skills within a linear chain of components. To enable skill path extracting from circuit graphs, we propose a three-step framework, consisting of decomposition, pruning, and post-pruning causal mediation. In particular, we offer a complete linear decomposition of the transformer model which leads to a disentangled computation graph. After pruning, we further adopt causal analysis techniques, including counterfactuals and interventions, to extract the final skill paths from the circuit graph. To underscore the significance of skill paths, we investigate three generic language skills-Previous Token Skill, Induction Skill, and In-Context Learning Skill-using our framework. Experiments support two crucial properties of these skills, namely stratification and inclusiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。