揭示大模型微调中对齐失效与恢复的动态机制。
Alignment Dynamics in LLM Fine-Tuning

- 提出可计算的对齐分数,解析微调时对齐变化的两大作用力。
- 发现前期对齐易被后续微调逆转,且后验分布越窄越难维持。
- 预测并验证了重播增强效应:旧对齐会加速新对齐恢复。
尽管大语言模型通过监督微调和人类反馈强化学习实现强对齐,但后续微调常导致对齐脆弱。现有解释或归因于梯度几何,或视为输出分布偏移,但缺乏将参数空间学习动态与函数空间对齐行为统一的框架。本文引入可计算的对齐得分,并推导其在微调中的闭式更新,构建统一的对齐动力学框架。分析显示,对齐更新由两大竞争成分决定:受当前对齐状态与模型后验狭窄程度共同控制的‘反弹力’,以及由训练分布与对齐/非对齐完成结果后验一致性决定的‘驱动力’。该分解解释了为何先前对齐会被后期微调逆转,且后验越窄,逆转越强。此外,框架预测存在‘重播预激活效应’:先前对齐会在后验中留下隐式印记,重新暴露时显著放大驱动力,加速再对齐。我们在安全对齐、涌现误对齐与情感设置中验证上述预测,均观察到一致的对齐逆转与重暴露下的加速再对齐。安全对齐的控制实验进一步证实反弹强度依赖于后验狭窄程度。这些结果为大模型微调中对齐的破坏与重启提供了统一的动力学视角。
原文摘要 · Abstract (English)
Although Large Language Models (LLMs) achieve strong alignment through supervised fine-tuning and reinforcement learning from human feedback, the alignment is often fragile under subsequent fine-tuning. Existing explanations either attribute alignment fragility to gradient geometry or characterize it as a distributional shift in model outputs, yet few provide a unified account that bridges parameter-space learning dynamics with function-space alignment behavior during fine-tuning. In this work, we introduce a tractable alignment score and derive its closed-form update during fine-tuning, yielding a unified framework for alignment dynamics. Our analysis decomposes alignment updates into two competing components: a \textbf{\color{red!60!black} Rebound Force}, governed jointly by the current alignment state and the narrowness of model distribution, and a \textbf{\color{green!60!black} Driving Force}, determined by how the training distribution aligns with outcome-conditioned posteriors over aligned and non-aligned completions. This decomposition explains why prior alignment can be reversed by later fine-tuning and why narrower posterior structure strengthens such reversal. Moreover, our framework predicts a \textbf{Rehearsal Priming Effect}: prior alignment leaves a latent posterior imprint that amplifies the effective Driving Force upon re-exposure, leading to faster re-alignment. We validate these predictions across safety alignment, emergent misalignment, and sentiment settings, demonstrating consistent alignment reversal and accelerated re-alignment under re-exposure. In addition, controlled experiments in safety alignment confirm the predicted dependence of rebound strength on posterior narrowness. Together, these results provide a unified dynamical perspective on how alignment is disrupted and reactivated during LLM fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。