arXiv:2503.20995cs.CL2025-03AAAI被引 4

用熵值筛选安全规则,提升多头奖励模型的准确性。

ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models

  • 根据评分熵值自动降低不一致规则的权重
  • 在RewardBench上超越随机、均匀等基线方法
  • 无需训练、可解释,适合各类安全对齐任务

大语言模型的安全对齐常依赖人类反馈强化学习(RLHF),需人工标注偏好数据集。为克服整体质量评分困难,近年研究转向基于多条安全规则的细粒度评分。本文发现:评分熵高的规则,其区分人类偏好的准确率反而较低。基于此,提出ENCORE方法,通过惩罚高熵规则来构建多头奖励。理论上,这类规则在Bradley-Terry损失下权重趋近于零,自然支持其被抑制。实验表明,ENCORE在RewardBench安全任务上持续优于随机、均匀加权、单头Bradley-Terry及LLM作为裁判等强基线。该方法完全免训练,跨数据集通用且保持可解释性,是多属性奖励建模的实用高效方案。

原文摘要 · Abstract (English)

The safety alignment of large language models (LLMs) often relies on reinforcement learning from human feedback (RLHF), which requires human annotations to construct preference datasets. Given the challenge of assigning overall quality scores to data, recent works increasingly adopt fine-grained ratings based on multiple safety rules. In this paper, we discover a robust phenomenon: Rules with higher rating entropy tend to have lower accuracy in distinguishing human-preferred responses. Exploiting this insight, we propose ENCORE, a simple entropy-guided method to compose multi-head rewards by penalizing rules with high rating entropy. Theoretically, we show that such rules yield negligible weights under the Bradley-Terry loss during weight optimization, naturally justifying their penalization. Empirically, ENCORE consistently outperforms strong baselines, including random and uniform weighting, single-head Bradley-Terry, and LLM-as-a-judge, etc. on RewardBench safety tasks. Our method is completely training-free, generally applicable across datasets, and retains interpretability, making it a practical and effective approach for multi-attribute reward modeling.

安全对齐奖励建模熵控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。