用多评委模型更真实地模拟人类偏好,提升大模型对齐效果。
Approximating Human Preferences Using a Multi-Judge Learned System
- 通过多个带评分标准的评委模型聚合判断,模拟多样化人类偏好。
- 在人和大模型评委上均验证了方法的鲁棒性,减少偏见与不稳定性。
- 适合需要精准偏好建模的场景,如RLHF奖励模型与智能路由系统。
将基于大语言模型的评判者与人类偏好对齐是重大挑战,因其难以校准,常受评分标准敏感性、偏见和不稳定性影响。解决该问题可推动关键应用,如为强化学习中的人类反馈(RLHF)构建可靠奖励模型,以及实现根据用户查询智能选择最优模型的路由系统。本文提出一种框架,通过学习聚合多个基于评分标准的评判者输出,建模多样化的角色化偏好。我们对比了该方法与基线方案的表现,并通过案例研究评估其在人类及大模型评判者偏见下的鲁棒性。主要贡献包括:可在大规模下合成角色化偏好标签的方法,以及两种实现聚合器的方式:广义加性模型(GAM)与多层感知机(MLP)。
原文摘要 · Abstract (English)
Aligning LLM-based judges with human preferences is a significant challenge, as they are difficult to calibrate and often suffer from rubric sensitivity, bias, and instability. Overcoming this challenge advances key applications, such as creating reliable reward models for Reinforcement Learning from Human Feedback (RLHF) and building effective routing systems that select the best-suited model for a given user query. In this work, we propose a framework for modeling diverse, persona-based preferences by learning to aggregate outputs from multiple rubric-conditioned judges. We investigate the performance of this approach against naive baselines and assess its robustness through case studies on both human and LLM-judges biases. Our primary contributions include a persona-based method for synthesizing preference labels at scale and two distinct implementations of our aggregator: Generalized Additive Model (GAM) and a Multi-Layer Perceptron (MLP).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。