arXiv:2603.02232cs.LGcs.AI2026-03被引 4

用数学模型统一处理人类对模型的分级评价,提升奖励建模效果。

Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback

  • 将李克特量表反馈视为有序回归问题,构建理论框架。
  • 学习阈值参数,比固定边距方法在多个任务上表现更优。
  • 适合需要精细人类反馈的对齐研究,如安全与推理任务。

奖励建模对于对齐大语言模型与人类偏好至关重要,但现有方法缺乏对等级偏好数据(如李克特量表)的严格数学框架。当人工标注者使用‘显著更好’、‘更好’、‘稍好’、‘几乎无差别’等分级评价时,现有方法通常采用启发式手段(如边距项或缩放因子)修改二元偏好模型(如Bradley-Terry)的损失函数,但这些方法未明确建模等级反馈的生成机制。本文提出一种理论严谨的框架,将李克特量表偏好下的奖励建模建模为离散有序回归问题,并推导出两种损失函数:负对数似然损失和全阈值损失,二者均能学习自然反映等级结构的阈值参数。与需手动设定固定边距或权重的启发式方法不同,本方法在统一概率框架下直接从数据中学习这些参数。在多个基准测试上的实验表明,该有序回归方法在聊天、推理与安全等多个评估类别中,性能始终优于或媲美现有启发式方法。本工作首次为在奖励模型训练中引入李克特量表反馈提供了原则性数学框架,推动了从对二元偏好模型的随意修改,向更有效利用细粒度人类反馈的转变。

原文摘要 · Abstract (English)

Reward modeling is crucial for aligning large language models with human preferences, yet current approaches lack a principled mathematical framework for leveraging ordinal preference data. When human annotators provide graded preferences on a Likert scale (e.g., significantly better, better, slightly better, negligibly better), existing methods typically apply ad-hoc heuristics, such as margin terms or scaling factors, to loss functions derived from binary preference models like Bradley-Terry. These approaches lack an underlying mathematical model for how ordinal preference data is generated. We present a theoretically grounded framework that formulates reward modeling with Likert scale preferences as a discrete ordinal regression problem. We derive two loss functions from this formulation: a negative log-likelihood loss and an all-threshold loss, both of which learn threshold parameters that naturally capture the ordinal structure of preferences. Unlike existing heuristic methods that manually specify fixed margins or scaling weights, our approach learns these parameters directly from data within a coherent probabilistic framework. Experimental results on multiple benchmarks demonstrate that our ordinal regression approach consistently achieves competitive or superior performance compared to existing heuristic methods across diverse evaluation categories including chat, reasoning, and safety tasks. Our work provides the first principled mathematical framework for incorporating Likert scale preferences into reward model training, moving beyond ad-hoc modifications of binary preference models to enable more effective utilization of fine-grained human feedback.

奖励建模有序回归人类偏好李克特量表

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。