arXiv:2509.04501cs.CLcs.AI2025-09被引 1

详解大模型指令微调中的强化学习方法,助你快速掌握核心原理。

Understanding Reinforcement Learning for Model Training, and future directions with GRAPE

  • 从零推导关键算法,用简化符号讲清每一步逻辑
  • 对比多种方法在指令微调中的表现差异与适用场景
  • 提出新框架GRAPE,适合想深入探索强化学习的科研者

本文系统阐述大模型指令微调中的一系列强化学习算法:SFT、拒绝采样、REINFORCE、TRPO、PPO、GRPO和DPO。现有解释常依赖先验知识、细节缺失或过于泛化复杂。本文以大模型为焦点,使用简化明确的符号,逐步推导每个方法,消除歧义,提升理解直观性。通过聚焦大模型应用场景,避免冗余抽象,降低认知负担。在此基础上,梳理了最新研究进展,并提出GRAPE(广义相对优势策略进化)作为未来研究方向,推动强化学习在模型训练中的进一步发展。

原文摘要 · Abstract (English)

This paper provides a self-contained, from-scratch, exposition of key algorithms for instruction tuning of models: SFT, Rejection Sampling, REINFORCE, Trust Region Policy Optimization (TRPO), Proximal Policy Optimization (PPO), Group Relative Policy Optimization (GRPO), and Direct Preference Optimization (DPO). Explanations of these algorithms often assume prior knowledge, lack critical details, and/or are overly generalized and complex. Here, each method is discussed and developed step by step using simplified and explicit notation focused on LLMs, aiming to eliminate ambiguity and provide a clear and intuitive understanding of the concepts. By minimizing detours into the broader RL literature and connecting concepts to LLMs, we eliminate superfluous abstractions and reduce cognitive overhead. Following this exposition, we provide a literature review of new techniques and approaches beyond those detailed. Finally, new ideas for research and exploration in the form of GRAPE (Generalized Relative Advantage Policy Evolution) are presented.

强化学习指令微调大模型GRAPE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。