研究大模型强化学习人类反馈的可扩展性,发现资源投入越多收益越低。
Does RLHF Scale? Exploring the Impacts From Data, Model, and Method
- 分析模型规模、数据多样性和推理预算对RLHF的影响
- 数据量和多样性提升显著改善奖励模型性能
- 大模型从固定奖励模型中获益更少,需优化资源配置
本研究系统探究了大语言模型中基于人类反馈的强化学习(RLHF)的可扩展性。尽管RLHF被视为大模型后训练的重要步骤,其扩展潜力仍不明确。我们分析了框架中的关键组件——模型规模、数据构成和推理预算——对性能的影响。结果表明,增加数据多样性和数量可提升奖励模型表现,有助于策略模型更好地扩展;在策略训练中,每提示生成更多回复样本初期能提升性能,但很快趋于饱和;更大的奖励模型仅带来小幅增益。此外,固定奖励模型下,更大规模的策略模型获益更少。总体而言,RLHF的扩展效率低于预训练,额外计算资源带来的回报递减。基于此,我们提出在计算约束下优化RLHF性能的策略。
原文摘要 · Abstract (English)
This study explores the scaling properties of Reinforcement Learning from Human Feedback (RLHF) in Large Language Models (LLMs). Although RLHF is considered an important step in post-training of LLMs, its scaling potential is still largely unknown. We systematically analyze key components in the RLHF framework--model size, data composition, and inference budget--and their impacts on performance. Our findings show that increasing data diversity and volume improves reward model performance, helping process-supervision models scale better. For policy training, more response samples per prompt boost performance initially but quickly plateau. And larger reward models offer modest gains in policy training. In addition, larger policy models benefit less from RLHF with a fixed reward model. Overall, RLHF scales less efficiently than pretraining, with diminishing returns from additional computational resources. Based on these observations, we propose strategies to optimize RLHF performance within computational limits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。