arXiv:2603.13985cs.AIcs.CL2026-03被引 4

对比SFT与RL在大模型微调中的优劣,揭示混合使用更高效。

Supervised Fine-Tuning versus Reinforcement Learning: A Study of Post-Training Methods for Large Language Models

  • 将监督微调与强化学习结合,构建统一训练框架。
  • 实证表明混合方法在任务准确率上优于单一策略。
  • 适合研究大模型后训练优化的开发者和研究人员。

预训练大语言模型具备广泛能力,但针对特定任务或领域提升准确性和可靠推理能力通常依赖于监督微调(SFT)或强化学习(RL)等后训练方法。尽管常被视为独立技术,近期理论与实证研究显示两者紧密关联。本文系统梳理SFT与RL的技术目标、算法结构与数据需求,深入分析其相互作用,提出融合二者优势的混合训练范式与集成框架。基于2023至2025年代表性应用研究,识别出后训练向混合模式快速演进的趋势,并总结各方法适用场景与效果差异。通过整合理论、方法与实证,本研究建立统一理解,为可扩展、高效、通用的大模型后训练指明未来方向。

原文摘要 · Abstract (English)

Pre-trained Large Language Model (LLM) exhibits broad capabilities, yet, for specific tasks or domains their attainment of higher accuracy and more reliable reasoning generally depends on post-training through Supervised Fine-Tuning (SFT) or Reinforcement Learning (RL). Although often treated as distinct methodologies, recent theoretical and empirical developments demonstrate that SFT and RL are closely connected. This study presents a comprehensive and unified perspective on LLM post-training with SFT and RL. We first provide an in-depth overview of both techniques, examining their objectives, algorithmic structures, and data requirements. We then systematically analyze their interplay, highlighting frameworks that integrate SFT and RL, hybrid training pipelines, and methods that leverage their complementary strengths. Drawing on a representative set of recent application studies from 2023 to 2025, we identify emerging trends, characterize the rapid shift toward hybrid post-training paradigms, and distill key takeaways that clarify when and why each method is most effective. By synthesizing theoretical insights, practical methodologies, and empirical evidence, this study establishes a coherent understanding of SFT and RL within a unified framework and outlines promising directions for future research in scalable, efficient, and generalizable LLM post-training.

大模型微调监督学习强化学习混合训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。