根据反馈数据动态选规则,让大模型更安全。
Data-adaptive Safety Rules for Training Reward Models
- 用最大差异法自动筛选最有效的评价规则
- 80亿参数模型在安全评测中排名第一
- 适合追求高安全性的模型训练者
强化学习从人类反馈(RLHF)常用于使大语言模型(LLMs)符合人类偏好,尤其是提升输出安全性。传统方法依赖成对响应的选择,但因人类意见差异及直接比较困难,近年趋向采用多维度指标或规则进行细粒度标注。关键挑战在于如何高效选择并应用这些规则以应对多样化的偏好数据。本文提出一种动态方法,针对每对响应自适应选择最重要规则。我们构建数学框架,利用成对响应间的最大差异,并理论证明该方法可最大化规则标注与真实偏好之间的互信息。随后,我们使用该自适应标注数据集训练了一个8B奖励模型,并在RewardBench上评估其效果。截至2025年1月25日,该模型在排行榜上安全性能领先,超越多个更大规模模型。
原文摘要 · Abstract (English)
Reinforcement Learning from Human Feedback (RLHF) is commonly employed to tailor models to human preferences, especially to improve the safety of outputs from large language models (LLMs). Traditionally, this method depends on selecting preferred responses from pairs. However, due to the variability in human opinions and the challenges in directly comparing two responses, there is an increasing trend towards fine-grained annotation approaches that evaluate responses using multiple targeted metrics or rules. The challenge lies in efficiently choosing and applying these rules to handle the diverse range of preference data. In this paper, we propose a dynamic method that adaptively selects the most important rules for each response pair. We introduce a mathematical framework that utilizes the maximum discrepancy across paired responses and demonstrate theoretically that this approach maximizes the mutual information between the rule-based annotations and the underlying true preferences. We then train an 8B reward model using this adaptively labeled preference dataset and assess its efficacy using RewardBench. As of January 25, 2025, our model achieved the highest safety performance on the leaderboard, surpassing various larger models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。