模型与任务对齐程度决定强化学习效果,强对齐时反直觉结果出现。
Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions
- 通过模型-任务对齐度判断反直觉RL现象是否成立
- 强对齐时单样本训练可媲美全数据集,负样本训练也有效
- 适用于已对齐模型,不适用于复杂任务场景
近期将强化学习(RL)应用于大语言模型(LLMs)取得显著进展,涌现出一系列令人惊讶但常违背直觉的现象:例如单个训练样本即可达到全数据集训练性能,奖励信号无需精确,仅用负样本训练也能媲美甚至超越复杂奖励方法。然而这些现象的适用条件及其失效边界仍不明确。本文识别出关键因素:预训练模型在目标任务上的对齐程度,以pass@k准确率衡量。通过跨模型架构和任务域的系统实验验证,发现尽管标准RL在各类场景下始终稳健,但多数反直觉现象仅在模型-任务对齐较强时出现;而在对齐较弱的困难任务中,这些技巧失效,标准RL仍有效。
原文摘要 · Abstract (English)
Recent advances in applying reinforcement learning (RL) to large language models (LLMs) have led to substantial progress. In particular, a series of remarkable yet often counterintuitive phenomena have been reported in LLMs, exhibiting patterns not typically observed in traditional RL settings. For example, notable claims include that a single training example can match the performance achieved with an entire dataset, that the reward signal does not need to be very accurate, and that training solely with negative samples can match or even surpass sophisticated reward-based methods. However, the precise conditions under which these observations hold - and, critically, when they fail - remain unclear. In this work, we identify a key factor that differentiates RL observations: whether the pretrained model already exhibits strong Model-Task Alignment, as measured by pass@k accuracy on the evaluated task. Through a systematic and comprehensive examination of a series of counterintuitive claims, supported by rigorous experimental validation across different model architectures and task domains, our findings show that while standard RL training remains consistently robust across settings, many of these counterintuitive results arise only when the model and task already exhibit strong model-task alignment. In contrast, these techniques fail to drive substantial learning in more challenging regimes, where standard RL methods remain effective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。