arXiv:2502.06533cs.CLcs.LG2025-02NAACL被引 36

通过放宽关键词的约束,提升大模型强化学习微调时的探索效率。

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

  • 在强化学习微调中动态调整关键词的KL惩罚,鼓励对重要词汇的探索。
  • 实验表明,该方法在算术任务上显著提升长程目标达成率。
  • 适合研究大模型探索机制或优化强化学习微调策略的研究者。

当前大语言模型实现长期目标的能力仍面临挑战。为解决此问题,可通过强化学习(RL)对预训练大模型进行微调,以探索优化目标的解决方案。然而,大模型的探索过程困难,需在发现新解与保持预训练模型特性之间取得平衡,以免损害基础能力,通常通过引入Kullback-Leibler(KL)惩罚来控制。本文研究了小语言模型在简单算术任务上的探索动态,揭示了不同预训练程度对探索行为的影响,并识别出对最终结果具有决定性影响的“关键令牌”。基于此,我们提出一种简单的KL惩罚修改策略,优先鼓励在关键令牌上的探索,从而显著提升强化学习微调阶段的效率。

原文摘要 · Abstract (English)

The ability to achieve long-term goals is a key challenge in the current development of large language models (LLMs). To address this, pre-trained LLMs can be fine-tuned with reinforcement learning (RL) to explore solutions that optimize a given goal. However, exploration with LLMs is difficult, as a balance has to be struck between discovering new solutions and staying close enough to the pre-trained model, so as not to degrade basic capabilities. This is typically controlled with a Kullback-Leibler (KL) penalty. In this paper, we investigate the exploration dynamics of a small language model on a simple arithmetic task. We show how varying degrees of pre-training influence exploration and demonstrate the importance of "critical tokens" which have a dramatic impact on the final outcome. Consequently, we introduce a simple modification to the KL penalty that favors exploration on critical tokens, increasing the efficiency of the RL fine-tuning stage.

强化学习大模型探索优化微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。