arXiv:2605.06076cs.CL2026-05被引 2

静态机制定位会滞后,无法有效指导大模型训练中的参数更新。

Navigating by Old Maps: The Pitfalls of Static Mechanistic Localization in LLM Post-Training

  • 用三个新指标追踪Transformer电路演化,发现参数更新时电路会自发迁移。
  • 实验表明,当前静态机制存在时间延迟,难以指导未来动态更新。
  • 提出需具备前瞻性的机制定位框架,适合关注模型可解释性研究者。

大型语言模型后训练中广泛采用的“定位-更新”范式,依赖于当前静态参数提取的机制进行精准参数调整。然而这一范式建立在未验证的根本假设之上:当前机制能否可靠指导未来的动态参数更新?本文系统追踪了监督微调(SFT)过程中Transformer电路的结构演化,揭示任务机制的内在动态。提出三类新度量:电路距离、电路稳定性和电路冲突,从神经迁移、语义稳定性与跨任务干扰三个维度分析演化规律。实证结果表明,电路在参数更新中存在固有的“自由演化”现象。因此,基于当前状态提取的静态机制必然存在时间延迟,从根本上不适用于指导未来状态。此外,通过解构现有方法的“有效性幻觉”,本文强调机制定位需具备前瞻性,并为后续研究提出预测性框架。

原文摘要 · Abstract (English)

The "Locate-then-Update" paradigm has become a predominant approach in the post-training of large language models (LLMs), identifying critical components via mechanistic interpretability for targeted parameter updates. However, this paradigm rests on a fundamental yet unverified assumption: can mechanisms derived from current static parameters reliably guide future dynamic parameter updates? To investigate this, we systematically track the structural evolution of Transformer circuits throughout the supervised fine-tuning (SFT) process, revealing the underlying dynamics of task mechanisms. We introduce three novel metrics-Circuit Distance, Circuit Stability, and Circuit Conflict-to analyze circuit evolution across three dimensions: neural migration, semantic stability, and cross-task interference. Our empirical results reveal that circuits inherently exhibit "Free Evolution" during parameter updates. Consequently, static mechanisms extracted from current states inevitably suffer from temporal latency, making them fundamentally inadequate for guiding future states. Moreover, by deconstructing the "illusion of effectiveness" in existing methods, this work underscores the necessity of "foresight" in mechanistic localization and proposes a predictive framework for future research.

机制解释模型演化大模型训练可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。