arXiv:2410.05584cs.LGcs.AI2024-10ICLR被引 25

发现奖励模型准确率不能可靠预测下游表现,可能误导评估。

Rethinking Reward Model Evaluation: Are We Barking up the Wrong Tree?

  • 用合成实验检验奖励模型准确率与策略性能的关系
  • 相同准确率下策略表现差异大,相关性很弱
  • 准确率测量方式影响预测能力,易引发过优化陷阱

奖励模型(RMs)在对齐语言模型与人类偏好中起关键作用。当前评估RMs主要依赖人工标注偏好数据集上的准确率,尽管方法简单且广泛使用,但其与下游策略性能的关系尚未充分探索。本文在合成设置下开展实验,研究不同准确率的RMs如何影响优化策略的表现。结果表明,准确率与下游性能仅有微弱正相关,且相似准确率的模型所生成的策略性能差异显著。此外,准确率的衡量方式显著影响其预测最终策略性能的能力。基于回归性古德哈特效应的视角,我们发现以准确率衡量RM质量时,难以全面捕捉其潜在的过优化问题。这揭示了仅依赖准确率评估RM效果的局限性。

原文摘要 · Abstract (English)

Reward Models (RMs) are crucial for aligning language models with human preferences. Currently, the evaluation of RMs depends on measuring accuracy against a validation set of manually annotated preference data. Although this method is straightforward and widely adopted, the relationship between RM accuracy and downstream policy performance remains under-explored. In this work, we conduct experiments in a synthetic setting to investigate how differences in RM measured by accuracy translate into gaps in optimized policy performance. Our findings reveal that while there is a weak positive correlation between accuracy and downstream performance, policies optimized towards RMs with similar accuracy can exhibit quite different performance. Moreover, we discover that the way of measuring accuracy significantly impacts its ability to predict the final policy performance. Through the lens of the Regressional Goodhart effect, we recognize that accuracy, when used for measuring RM quality, can fail to fully capture the potential RM overoptimization. This underscores the inadequacy of relying solely on accuracy to reflect their impact on policy optimization.

奖励模型评估方法过优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。