arXiv:2504.02976cs.LGcs.AI2025-04被引 2

发现模型中定义性知识集中存储,关联性知识则分散分布。

Localized Definitions and Distributed Reasoning: A Proof-of-Concept Mechanistic Interpretability Study via Activation Patching

  • 用激活修补法定位关键神经层,分析知识存储位置。
  • 首层修复恢复56%正确偏好,末层修复实现100%准确率。
  • 适合关注模型可解释性与精准编辑的研究者。

本研究通过因果层归因激活修补法(CLAP),探究微调后的GPT-2模型在9,958篇PubMed摘要(癫痫:20,595次提及,脑电图:11,674次提及,发作:13,921次提及)上的知识表征定位问题。方法包括缓存正确与错误激活、计算逻辑差异,并修补错误激活以评估恢复效果。结果表明:首先,修补第一前馈层可恢复56%的正确偏好,说明关联知识分布于多层;其次,修补最终输出层实现100%准确恢复,且定义类问题的清洁逻辑差异更强,表明定义性知识局部化;第三,卷积层修补仅恢复13.6%,说明低层特征对高层推理贡献极小。统计分析确认层间效应显著(p<0.01)。研究揭示事实性知识更局部化,而关联性知识依赖分布式表示,且编辑效果取决于任务类型,为模型可解释性与任务自适应更新提供依据。

原文摘要 · Abstract (English)

This study investigates the localization of knowledge representation in fine-tuned GPT-2 models using Causal Layer Attribution via Activation Patching (CLAP), a method that identifies critical neural layers responsible for correct answer generation. The model was fine-tuned on 9,958 PubMed abstracts (epilepsy: 20,595 mentions, EEG: 11,674 mentions, seizure: 13,921 mentions) using two configurations with validation loss monitoring for early stopping. CLAP involved (1) caching clean (correct answer) and corrupted (incorrect answer) activations, (2) computing logit difference to quantify model preference, and (3) patching corrupted activations with clean ones to assess recovery. Results revealed three findings: First, patching the first feedforward layer recovered 56% of correct preference, demonstrating that associative knowledge is distributed across multiple layers. Second, patching the final output layer completely restored accuracy (100% recovery), indicating that definitional knowledge is localised. The stronger clean logit difference for definitional questions further supports this localized representation. Third, minimal recovery from convolutional layer patching (13.6%) suggests low-level features contribute marginally to high-level reasoning. Statistical analysis confirmed significant layer-specific effects (p<0.01). These findings demonstrate that factual knowledge is more localized and associative knowledge depends on distributed representations. We also showed that editing efficacy depends on task type. Our findings not only reconcile conflicting observations about localization in model editing but also emphasize on using task-adaptive techniques for reliable, interpretable updates.

可解释性知识定位模型编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。