arXiv:2506.21560cs.CLcs.AI2025-06

用强化学习微调小模型,提升指令遵循与数学推理能力

Reinforcement learning fine-tuning of language model for instruction following and math reasoning

  • 采用RLOO与DeBERTa奖励模型实现最佳对齐效果
  • 结合合成数据与外部验证器,数学推理准确率显著提升
  • 适合资源有限下训练高效精准的小模型开发者参考

本研究探讨了强化学习(RL)微调技术在小型语言模型(Qwen2.5-0.5B Base)上的应用,针对指令遵循与数学推理两大挑战任务。对比监督微调(SFT)、基于偏好数据的直接偏好优化(DPO)以及使用奖励模型的Reinforce Leave-One-Out(RLOO)方法,实验表明:采用DeBERTa奖励模型的RLOO实现最优对齐;DPO表现稳定且优秀。在数学推理任务中,通过合成数据增强与外部验证器配合的best-of-N采样,显著提升准确率,展现出微调与推理阶段工具结合的巨大潜力。研究揭示了轻量级、任务对齐小模型训练中的关键权衡与实用策略。

原文摘要 · Abstract (English)

This study investigates the effectiveness of reinforcement learning (RL) fine-tuning techniques on a compact language model (Qwen2.5-0.5B Base) for two challenging tasks: instruction following and mathematical reasoning. We compare supervised fine-tuning (SFT), Direct Preference Optimization (DPO) using preference-labeled data, and Reinforce Leave-One-Out (RLOO) with reward models. Our experiments show that RLOO with DeBERTa reward modeling achieves the best alignment, while DPO provides strong and consistent results. For math reasoing tasks, synthetic data augmentation and best-of-N sampling with an external verifier significantly improve accuracy, showing the potential of combining fine-tuning with inference-time tools. This study highlights key trade-offs and practical strategies for training lightweight, task-aligned small-scale language models.

强化学习小模型数学推理指令遵循

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。