arXiv:2512.25023cs.LG2025-12NeurIPS被引 1

通过局部相对信号提升奖励模型的样本效率与鲁棒性

ResponseRank: Data-Efficient Reward Modeling through Preference Strength Learning

  • 利用局部相对代理信号(如响应时间)推断偏好强度,避免全局噪声干扰
  • 在合成数据、语言建模和强化学习控制任务中均实现更高样本效率
  • 提出新指标PDC,分离度量效用学习与排序准确性的差异

二元选择常用于人类反馈强化学习(RLHF),但仅提供偏好的方向信息。例如,某人可能更偏好苹果而非橙子,也更偏好香蕉而非葡萄,但哪种偏好更强?偏好强度对不确定条件下的决策和模型泛化至关重要,但难以可靠测量。响应时间、标注者一致性等元数据可作为强度代理,但通常存在噪声且易受混杂因素影响。本文提出ResponseRank方法,通过局部构建分层结构,仅比较同一分层内的相对代理信号,以更稳健地推断响应的偏好强度。该方法能有效利用噪声强度信号,同时最小化对信号分布的假设。主要贡献包括:(1) 提出ResponseRank,一种基于局部有效相对强度信号的偏好强度学习新方法;(2) 在合成偏好学习(模拟响应时间)、语言建模(标注者一致性)和强化学习控制任务(模拟回合回报)中验证了其更高的样本效率与鲁棒性;(3) 提出皮尔逊距离相关性(PDC)作为新评估指标,可独立衡量基数效用学习与序数准确性。

原文摘要 · Abstract (English)

Binary choices, as often used for reinforcement learning from human feedback (RLHF), convey only the direction of a preference. A person may choose apples over oranges and bananas over grapes, but which preference is stronger? Strength is crucial for decision-making under uncertainty and generalization of preference models, but hard to measure reliably. Metadata such as response times and inter-annotator agreement can serve as proxies for strength, but are often noisy and confounded. We propose ResponseRank to address the challenge of learning from noisy strength signals. Our method uses relative differences in proxy signals to rank responses to pairwise comparisons by their inferred preference strength. To control for systemic variation, we compare signals only locally within carefully constructed strata. This enables robust learning of utility differences consistent with strength-derived rankings while making minimal assumptions about the strength signal. Our contributions are threefold: (1) ResponseRank, a novel method that robustly learns preference strength by leveraging locally valid relative strength signals; (2) empirical evidence of improved sample efficiency and robustness across diverse tasks: synthetic preference learning (with simulated response times), language modeling (with annotator agreement), and RL control tasks (with simulated episode returns); and (3) the Pearson Distance Correlation (PDC), a novel metric that isolates cardinal utility learning from ordinal accuracy.

奖励建模偏好学习样本效率强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。