arXiv:2608.24482cs.LGcs.AI2026-08

用预训练参数预测微调后关键神经元,提升模型优化效率

Beyond Static Interpretability: Anticipating Post-SFT Mechanisms from Pre-SFT Parameters for Better Tuning

论文配图:Beyond Static Interpretability: Anticipating Post-SFT Mechanisms from Pre-SFT Parameters for Better Tuning
图 1 · 摘自论文原文
  • 基于预训练参数和目标数据,预测微调后的关键神经元位置
  • 在多个模型规模下实现更优的参数高效微调效果
  • 突破传统可解释性滞后问题,适合追求高效微调的研究者

机制定位将机制可解释性与训练后优化相连接,通过可解释方法定位关键参数,并以‘定位-微调’范式指导参数高效的监督微调(SFT)。然而,由于机制可解释性的回顾性,直接对预-SFT模型进行解释会得出误导性结论。尤其在新任务中,初始识别出的神经元与最终模型所依赖的神经元差异巨大,引入偏差并干扰SFT。为此,我们提出一种前瞻性的定位框架,仅使用预-SFT参数和目标数据,准确估计微调后的可解释状态。理论上,我们将SFT建模为连续参数演化,利用泰勒展开严格关联微调后机制目标与预训练模型的动态梯度;实践中,设计了神经元级与组件级双粒度定位流程。大量实验表明,该方法不仅提供更优的SFT引导,还在模型规模增大时展现出鲁棒性与时间可扩展性。本工作突破传统可解释性根本局限——无法在训练前识别任务关键机制——开创性地将机制可解释性与定向优化结合,推动可解释性迈向预测前沿。

原文摘要 · Abstract (English)

Mechanistic Localization bridges mechanistic interpretability and post-training optimization by isolating critical parameters via interpretative approaches and then guiding parameter-efficient Supervised Fine-Tuning (SFT) in a ``locating-then-tuning'' paradigm. However, due to the retrospective nature of mechanistic interpretability, directly interpreting pre-SFT models introduces misleading conclusions. Specifically for novel tasks, initially identified neurons differ drastically from those governing the final model, introducing biases that actively disrupt SFT. To address this, we propose a forward-looking localization framework that accurately estimates the post-SFT interpretability state using only pre-SFT parameters and the target dataset. Theoretically, we model SFT as a continuous parameter evolution, leveraging Taylor expansion to rigorously bridge the post-tuning mechanistic objective with the pre-SFT model's dynamic gradients. Practically, we design dual-granularity (neuron- and component-level) localization pipelines. Extensive experiments demonstrate that our approach not only provides superior SFT guidance but also exhibits robust performance and temporal scalability across increasing model sizes. This work transcends the fundamental limitation of traditional interpretability-its inability to identify task-critical mechanisms before they are trained-pioneering a predictive frontier that unites mechanistic interpretability with targeted optimization.

可解释性微调优化机制定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。