通过行为注入提升大模型对强化学习的响应能力
Behavior Injection: Preparing Language Models for Reinforcement Learning
- 在微调数据中注入探索与利用行为,增强模型对强化学习的适应性
- 实验证明,经处理的模型在推理基准上强化学习增益显著提升
- 适合希望稳定提升大模型推理能力的研究者和工程师
强化学习(RL)已成为激励大型语言模型(LLMs)推理能力的强大后训练技术。然而,LLMs 对 RL 微调的响应极不一致:部分模型性能显著提升,而另一些则停滞甚至下降。为理解这一差异,我们分析了每步强化学习目标的影响,识别出有效后训练的两个关键条件:(1) 强化学习信息丰富的轨迹准确率,(2) 强大的数据共影响性,即训练数据对其他样本性能的影响程度。基于这些洞察,我们提出行为注入,一种任务无关的数据增强方法,在进行强化学习前应用于监督微调(SFT)数据。该方法通过引入探索性和利用性行为,丰富了 SFT 数据,使模型更具备强化学习就绪性。我们在两个推理基准上使用多个基础模型评估该方法,结果表明,这种理论驱动的数据增强可显著提升强化学习带来的性能增益,相较于预强化学习模型表现更优。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has emerged as a powerful post-training technique to incentivize the reasoning ability of large language models (LLMs). However, LLMs can respond very inconsistently to RL finetuning: some show substantial performance gains, while others plateau or even degrade. To understand this divergence, we analyze the per-step influence of the RL objective and identify two key conditions for effective post-training: (1) RL-informative rollout accuracy, and (2) strong data co-influence, which quantifies how much the training data affects performance on other samples. Guided by these insights, we propose behavior injection, a task-agnostic data augmentation scheme applied prior to RL. Behavior injection enriches the supervised finetuning (SFT) data by seeding exploratory and exploitative behaviors, effectively making the model more RL-ready. We evaluate our method across two reasoning benchmarks with multiple base models. The results demonstrate that our theoretically motivated augmentation can significantly increase the performance gain from RL over the pre-RL model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。