arXiv:2605.16615cs.LG2026-05

提出一种无需强假设的评估偏好学习方法,更真实地捕捉人类与大模型的评价逻辑。

Learning What Evaluators Value: A Reliable Approach to Modeling Evaluator Preferences

论文配图:Learning What Evaluators Value: A Reliable Approach to Modeling Evaluator Preferences
图 1 · 摘自论文原文
  • 仅假设评价函数各维度非递减,避免传统方法对线性等强假设的依赖
  • 理论证明在线性假设下可无损学习任意偏好函数,且实验证明鲁棒性强
  • 适用于理解评审偏好、提升学术同行评审公平性等实际场景

在诸多应用场景中,人类或大语言模型评估者会基于多个相关标准对项目或个体进行综合评估。例如,在招生中,委员会依据考试成绩、GPA和科研经历等评估候选人整体匹配度;在医疗中,临床医生根据患者症状报告进行初步诊断与风险评估。这些场景均涉及从多维标准映射到总体评价的过程,反映了评估者的潜在偏好。本文聚焦于学习这些偏好这一核心问题。现有方法常依赖特定偏好建模假设,可能在真实世界中严重偏离。本文仅做最小化假设:偏好函数在每个坐标上非递减,该假设在多数评估场景中合理。我们理论上分析了多种常见假设下的模型误配程度,表明其可能导致学习偏差及下游任务性能下降。为此,我们提出一种对模型误配鲁棒的偏好学习算法,并理论证明:当线性假设成立时,该算法能无损学习任意偏好函数。通过合成模拟与真实数据验证,结果表明该算法能可靠学习偏好,并揭示大模型与人类偏好的关键特征。最后,我们以同行评审为例,展示如何利用该方法洞察审稿人偏好并提升评审公平性。

原文摘要 · Abstract (English)

In many applications, human and LLM evaluators use assessments of relevant criteria to create an overall evaluation for an item or individual. For example, in admissions, committees assess candidates on attributes such as test scores, GPA, and research experience to evaluate their overall fit for the program. Another example arises in medical care, where clinicians use patient reports of symptoms to consider preliminary diagnoses and assess risks. Each setting involves mapping multiple criteria to an overall evaluation---a process that reflects the evaluator's underlying preferences. We focus on the fundamental question of learning these preferences. Many applications of this problem make specific modeling assumptions on evaluator preferences that may be substantially violated in the real world. We make the minimal assumption that the preference function is coordinate-wise non-decreasing, which is reasonable in a large number of evaluation settings. We theoretically characterize the severity of model mismatch for many common assumptions and show that it can lead to significant issues for learning evaluator preferences and other important downstream tasks. We then present an algorithm for learning evaluators' preferences that is robust to model mismatch. We prove theoretically that our algorithm can learn any preference function without sacrificing performance when the linearity assumption holds. Evaluations of our algorithm with synthetic simulations and real-world data confirm its ability to learn preferences robustly and illustrate key aspects of LLM and human preferences. To conclude, we present a case study of how to use our method to provide insights into reviewer preferences and increase fairness in the peer review process.

偏好学习评估模型公平性同行评审

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。