arXiv:2606.28525cs.LGcs.AI2026-06

发现模型微调后行为会回退,提出引力解释机制。

A Gravitational Interpretation of Fine-Tuning Reversion

论文配图:A Gravitational Interpretation of Fine-Tuning Reversion
图 1 · 摘自论文原文
  • 用几何视角看训练历史,早期大阶段形成主导行为流形。
  • 微调时出现指向历史方向的回退分量,对齐度从0.429升至0.647。
  • 阻断该方向可降低有害性至8.5%,适合关注安全性的研究者。

在无害数据上进行微调可能部分逆转早期训练中习得的行为。安全性能在看似良性的对齐更新后退化,被遗忘的能力可能重现,潜在特征可通过看似无关的监督传递,且其他生成场景中也存在类似对齐脆弱性。我们认为这些现象可通过共同的训练历史视角理解。我们的假设是几何性的:早期大规模训练阶段形成主导行为流形,后续对齐或专业化阶段只是其上的浅层位移。因此,后续微调会继承一个指向主导流形见证的持久回退分量。我们称之为微调回退的引力解释。在主要设置中,表征漂移迅速获得沿历史定义回退方向(v_rev)的分量。主实验中,与v_rev的对齐度从首次更新后的0.429 ± 0.052提升至第20步的0.647 ± 0.021。在24组运行-步骤对中,所有观测到的对齐度均超过各向同性激活空间零模型的p99。我们证明,选择性阻断沿v_rev的运动可使T=100时的对齐度从0.648 ± 0.009降至-0.211 ± 0.021,并将有害性从19.0% ± 4.0%降至8.5% ± 1.5%,任务损失极小。这些结果支持v_rev是早期对齐后回退动力学中的因果相关中介。重要的是,我们不认为v_rev是唯一的安全方向,也不声称主导流形可直接观测;而是识别出一个稳健、由历史定义的方向,能解释并部分控制早期回退动态。

原文摘要 · Abstract (English)

Fine-tuning on harmless data can partially undo behaviors acquired earlier in training. Safety can erode under benign post-alignment updates, unlearned capabilities can re-emerge, latent traits can transfer through apparently unrelated supervision, and related post-alignment fragility appears in other generative settings. We argue these phenomena are usefully viewed through a common training-history lens. Our hypothesis is geometric: large early training phases create dominant behavioral manifolds, while later alignment or specialization phases are shallower displacements from them. Subsequent fine-tuning can therefore inherit a persistent reversion component pointing back toward a witness of the dominant manifold. We call this the gravitational interpretation of fine-tuning reversion. Across our main settings, representational drift rapidly acquires a component along a history-defined reversion direction (v_rev). In our main track, alignment with v_rev rises from cos = 0.429 +/- 0.052 after the first update to 0.647 +/- 0.021 by step 20. Across 24 run-step pairs, every observed alignment exceeds the p99 of an isotropic activation-space null. We demonstrate that selectively blocking motion along v_rev changes the final alignment at T=100 from 0.648 +/- 0.009 to -0.211 +/- 0.021 and reduces harmfulness from 19.0% +/- 4.0% to 8.5% +/- 1.5% with little task cost. These results support v_rev as a causally relevant mediator of early post-alignment reversion in our setup. Importantly, we do not claim that v_rev is the unique safety direction, nor that the dominant manifold is directly observed; rather, we identify a robust, history-defined direction that explains and partially controls early reversion dynamics.

模型安全微调回退训练历史几何解释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。