用动态对抗测试评估大模型奖励模型泛化能力
Evaluating Reward Model Generalization via Pairwise Maximum Discrepancy Competitions
- 通过对比两模型分歧最大样本,自动筛选高争议测试用例
- 10个奖励模型在新评测下排名大幅变化,揭示传统评估偏差
- 适合研究奖励模型泛化性与评估方法的学者
奖励模型(RMs)在对齐大语言模型中起关键作用,但其实际效果取决于对未见提示和分布漂移的泛化能力。现有评估多依赖静态标注偏好数据集,覆盖有限且难以真实反映开放世界中的泛化表现。本文提出配对最大差异竞赛(PMDC),一种基于大规模无标注开放域提示池的动态、高效评估框架。该方法主动选取两模型分歧最大的提示-响应对,生成紧凑的高争议测试集。这些案例由人工裁判判定,结果通过Bradley-Terry模型聚合,获得各模型的全局排名与成对胜率图谱。我们用PMDC重新评估了10个代表性奖励模型,发现其排名与传统基准相比出现显著重塑。定性分析进一步揭示系统性泛化失败模式,为改进奖励建模提供重要洞见。
原文摘要 · Abstract (English)
Reward models (RMs) are central to aligning large language models, yet their practical effectiveness hinges on generalization to unseen prompts and shifting distributions. Most existing RM evaluations rely on static, pre-annotated preference datasets, which provide limited coverage and often fail to faithfully assess generalization in open-world settings. We introduce Pairwise Maximum Discrepancy Competition (PMDC), a dynamic and annotation-efficient framework for evaluating RM generalization using a large, unlabeled, open-domain prompt pool. PMDC actively selects prompt--response pairs that maximize disagreement between two RMs, yielding a compact set of highly contentious test cases. These cases are adjudicated by an oracle, and the resulting outcomes are aggregated via a Bradley--Terry model to produce a global ranking and pairwise win-rate landscape of RMs. We apply PMDC to re-evaluate 10 representative RMs and observe substantial rank reshuffling compared with conventional benchmarks. Qualitative analyses further uncover systematic generalization failures, providing valuable insights for improving reward modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。