通过可控探索与精修离线融合,提升小模型推理能力
Recall-Extend Dynamics: Enhancing Small Language Models through Controlled Exploration and Refined Offline Integration
- 设计动态调控机制,平衡离线数据与在线强化学习
- 通过熵变化比率调节离线微调权重,缓解小模型探索不足
- 适合研究小模型增强与高效知识迁移的开发者
现有研究通过可验证奖励的强化学习(RLVR)显著提升了大语言模型的推理能力,但小语言模型(SLMs)的推理增强仍缺乏充分探索。将大模型蒸馏数据与小模型自身RLVR结合是自然思路,但仍面临诸多挑战。本文提出召回-扩展动态机制(RED),从探索空间变化角度出发,协调离线蒸馏与在线强化学习。针对离线数据中的插入问题,专门设计优化方案。通过监测模型在离线与在线数据上熵的变化比率,动态调节离线SFT权重,解决小模型探索空间不足及蒸馏过程冗余复杂的问题。此外,为缓解离线数据与当前策略间的分布差异,设计基于样本准确率的策略转移机制,动态选择模仿离线数据或自主学习。
原文摘要 · Abstract (English)
Many existing studies have achieved significant improvements in the reasoning capabilities of large language models (LLMs) through reinforcement learning with verifiable rewards (RLVR), while the enhancement of reasoning abilities in small language models (SLMs) has not yet been sufficiently explored. Combining distilled data from larger models with RLVR on small models themselves is a natural approach, but it still faces various challenges and issues. Therefore, we propose \textit{\underline{R}}ecall-\textit{\underline{E}}xtend \textit{\underline{D}}ynamics(RED): Enhancing Small Language Models through Controlled Exploration and Refined Offline Integration. In this paper, we explore the perspective of varying exploration spaces, balancing offline distillation with online reinforcement learning. Simultaneously, we specifically design and optimize for the insertion problem within offline data. By monitoring the ratio of entropy changes in the model concerning offline and online data, we regulate the weight of offline-SFT, thereby addressing the issues of insufficient exploration space in small models and the redundancy and complexity during the distillation process. Furthermore, to tackle the distribution discrepancies between offline data and the current policy, we design a sample-accuracy-based policy shift mechanism that dynamically chooses between imitating offline distilled data and learning from its own policy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。