用经济学原理优化语言模型奖励,让生成内容更有用且更安全。
Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models
- 基于效用理论设计奖励变换,增强对低分的敏感度。
- 相比线性加权平均,生成文本更助人且危害更少。
- 适合关注RLHF训练质量与安全性的研究者使用。
当前基于强化学习反馈训练大语言模型的方法通常采用多个奖励函数输出的线性平均。这种方法忽略了各奖励维度的特性及相互依赖关系,可能导致生成结果次优。本文揭示了线性聚合在某些情况下会引发不良生成属性,并提出一种受经济效用理论(特别是Inada条件)启发的奖励变换方法:该方法增强对低奖励值的敏感度,同时降低对高奖励值的响应。与传统的加权平均基线相比,采用Inada变换的反馈机制显著提升模型表现。定量与定性分析均表明,经此训练的模型生成内容更具帮助性且危害更低。
原文摘要 · Abstract (English)
Current methods that train large language models (LLMs) with reinforcement learning feedback, often resort to averaging outputs of multiple rewards functions during training. This overlooks crucial aspects of individual reward dimensions and inter-reward dependencies that can lead to sub-optimal outcomes in generations. In this work, we show how linear aggregation of rewards exhibits some vulnerabilities that can lead to undesired properties of generated text. We then propose a transformation of reward functions inspired by economic theory of utility functions (specifically Inada conditions), that enhances sensitivity to low reward values while diminishing sensitivity to already high values. We compare our approach to the existing baseline methods that linearly aggregate rewards and show how the Inada-inspired reward feedback is superior to traditional weighted averaging. We quantitatively and qualitatively analyse the difference in the methods, and see that models trained with Inada-transformations score as more helpful while being less harmful.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。