arXiv:2507.00018cs.LGcs.AI2025-07NeurIPS被引 21

揭示大模型微调中SFT与偏好学习的统一理论机制

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections

  • 发现SFT本质是隐式奖励学习,与DPO共享最优策略-奖励子空间
  • 提出学习率降低法,指令跟随任务性能提升25%相对增益
  • 拓展了对齐理论,为优化微调方法提供数学基础,适合研究者参考

后训练是将预训练语言模型适配真实任务的关键阶段,通过示范或偏好信号学习至关重要。本文构建了一个统一的理论框架,连接大语言模型后训练中的监督微调(SFT)与偏好学习。通过严谨的数学推导,我们证明SFT与直接偏好优化(DPO)等方法均作用于同一最优策略-奖励子空间,且SFT是隐式奖励学习的特例。分析揭示传统SFT存在关键缺陷:分布匹配中的KL散度在优化过程中对策略恒定,无法约束模型更新。为此,我们提出一种简单有效的学习率降低策略,在指令遵循任务中实现最高25%的相对性能提升和6%的绝对胜率增长。此外,我们从多种f-散度函数推导出保留KL项的替代SFT目标,进一步提升后DPO模型表现。最后,我们将偏好学习中关于模型logits与Q函数的关系理论扩展至SFT场景,给出数学推导与实验验证。

原文摘要 · Abstract (English)

Post-training processes are essential phases in grounding pre-trained language models to real-world tasks, with learning from demonstrations or preference signals playing a crucial role in this adaptation. We present a unified theoretical framework bridging Supervised Fine-Tuning (SFT) and preference learning in Large Language Model (LLM) post-training. Through rigorous mathematical derivation, we demonstrate that both SFT and preference learning methods like Direct Preference Optimization (DPO) operate within the same optimal policy-reward subspace, with SFT representing a special case of implicit reward learning. Our analysis reveals a critical limitation in conventional SFT: the KL divergence term in distribution matching becomes constant with respect to the policy during optimization, failing to constrain model updates. To address this, we propose a simple yet effective learning rate reduction approach that yields significant performance improvements (up to \textbf{25\%} relative gain and \textbf{6\%} absolute win rate increase in instruction following tasks. Additionally, we derive alternative SFT objectives from various f-divergence functions that preserve the KL term during optimization, further enhancing post-DPO model performance. Finally, we extend the theoretical relationship between LLM logits and Q-functions from preference learning to the SFT context, providing mathematical derivations and experimental validation.

大模型微调偏好学习强化学习理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。