选对基础模型能让奖励模型性能提升14%。
A Systematic Analysis of Base Model Choice for Reward Modeling
- 系统测试不同大模型作基础模型对奖励模型的影响。
- 最佳选择可使性能比默认方案高出14%。
- 用少量基准测试组合可提升模型选择效果18%。
基于人类反馈的强化学习(RLHF)及核心的奖励建模已成为训练强大大语言模型的关键环节。在训练高质量奖励模型(RMs)时,一个常被忽视的因素是基础模型的选择,而随着大模型数量激增,这一选择愈发困难。本文系统分析了基础模型选择对奖励建模性能的影响。结果表明,相比最常见的默认选择,性能最高可提升14%。我们还展示了现有某些基准测试与下游性能间的强统计关联性。此外,通过结合少量基准测试结果,可在前5-10名中实现平均18%的模型选择性能提升。最后,我们分析了不同后训练步骤对最终性能的影响,并探索使用估计数据分布以降低性能预测误差。
原文摘要 · Abstract (English)
Reinforcement learning from human feedback (RLHF) and, at its core, reward modeling have become a crucial part of training powerful large language models (LLMs). One commonly overlooked factor in training high-quality reward models (RMs) is the effect of the base model, which is becoming more challenging to choose given the rapidly growing pool of LLMs. In this work, we present a systematic analysis of the effect of base model selection on reward modeling performance. Our results show that the performance can be improved by up to 14% compared to the most common (i.e., default) choice. Moreover, we showcase the strong statistical relation between some existing benchmarks and downstream performances. We also demonstrate that the results from a small set of benchmarks could be combined to boost the model selection ($+$18% on average in the top 5-10). Lastly, we illustrate the impact of different post-training steps on the final performance and explore using estimated data distributions to reduce performance prediction error.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。