arXiv:2505.16984cs.LGcs.CL2025-05NeurIPS被引 56

统一监督与强化微调,提升大模型推理能力

UFT: Unifying Supervised and Reinforcement Fine-Tuning

  • 将监督与强化微调融合为单一训练流程
  • 在不同模型规模下均优于传统方法
  • 理论上突破强化微调的样本复杂度瓶颈

后训练显著提升了大语言模型的推理能力。主流方法可分为监督微调(SFT)和强化微调(RFT)。SFT效率高,适合小模型,但易过拟合,限制大模型推理;RFT泛化性好,但依赖基础模型质量。为此,我们提出统一微调(UFT),将SFT与RFT整合为统一流程。UFT使模型既能有效探索解空间,又能利用有信息量的监督信号,弥合了记忆与思考之间的差距。实验表明,UFT在各种模型规模下均优于SFT和RFT。此外,我们理论证明UFT打破了RFT固有的指数级样本复杂度瓶颈,首次表明统一训练可使长时程推理任务的收敛速度呈指数级加速。

原文摘要 · Abstract (English)

Post-training has demonstrated its importance in enhancing the reasoning capabilities of large language models (LLMs). The primary post-training methods can be categorized into supervised fine-tuning (SFT) and reinforcement fine-tuning (RFT). SFT is efficient and well-suited for small language models, but it may lead to overfitting and limit the reasoning abilities of larger models. In contrast, RFT generally yields better generalization but depends heavily on the strength of the base model. To address the limitations of SFT and RFT, we propose Unified Fine-Tuning (UFT), a novel post-training paradigm that unifies SFT and RFT into a single, integrated process. UFT enables the model to effectively explore solutions while incorporating informative supervision signals, bridging the gap between memorizing and thinking underlying existing methods. Notably, UFT outperforms both SFT and RFT in general, regardless of model sizes. Furthermore, we theoretically prove that UFT breaks RFT's inherent exponential sample complexity bottleneck, showing for the first time that unified training can exponentially accelerate convergence on long-horizon reasoning tasks.

大模型微调推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。