arXiv:2510.00263cs.CL2025-10被引 4

让AI裁判学会预测人类偏好分布,提升判断可信度。

Judging with Confidence: Calibrating Autoraters to Preference Distributions

  • 用概率分布替代单一标签训练AI裁判
  • 预测结果更贴近真实人群偏好,偏差更低
  • 适合需要可解释评估的AI对齐研究

大型语言模型(LLM)对齐人类价值观越来越多地依赖其他LLM作为自动化评判者(即「autoraters」)。然而,其可靠性受限于一个根本问题:它们在离散偏好标签上训练,强制将单一真实答案应用于本就主观、模糊或复杂的任务。本文主张,可靠的autorater应学习目标人群偏好的完整分布。我们提出一种通用框架,用于将概率型autorater校准到任意给定偏好分布。通过形式化问题,设计两种学习方法:1)针对密集概率标签的直接监督微调;2)针对稀疏二元标签的强化学习。实验证明,采用分布匹配目标微调后,生成的概率预测更贴合目标偏好分布,校准性更好,位置偏差显著降低,同时保持客观任务性能。

原文摘要 · Abstract (English)

The alignment of large language models (LLMs) with human values increasingly relies on using other LLMs as automated judges, or ``autoraters''. However, their reliability is limited by a foundational issue: they are trained on discrete preference labels, forcing a single ground truth onto tasks that are often subjective, ambiguous, or nuanced. We argue that a reliable autorater must learn to model the full distribution of preferences defined by a target population. In this paper, we propose a general framework for calibrating probabilistic autoraters to any given preference distribution. We formalize the problem and present two learning methods tailored to different data conditions: 1) a direct supervised fine-tuning for dense, probabilistic labels, and 2) a reinforcement learning approach for sparse, binary labels. Our empirical results show that finetuning autoraters with a distribution-matching objective leads to verbalized probability predictions that are better aligned with the target preference distribution, with improved calibration and significantly lower positional bias, all while preserving performance on objective tasks.

AI对齐自动评判偏好建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。