arXiv:2604.00200cs.LG2026-04

让AI在离线学习中兼顾性能与安全,用偏好反馈自动优化决策。

Offline Constrained RLHF with Multiple Preference Oracles

论文配图:Offline Constrained RLHF with Multiple Preference Oracles
图 1 · 摘自论文原文
  • 基于多源偏好数据,通过最大似然估计构建奖励模型。
  • 首次给出离线约束强化学习的有限样本性能保证。
  • 适合需要保障公平性或安全性的高风险场景应用。

我们研究了基于多个偏好信号的离线约束强化学习。针对需在性能与安全或公平之间权衡的应用,目标是在保证受保护群体福利不低于阈值的前提下,最大化整体群体效用。利用参考策略下收集的成对比较数据,通过最大似然方法估计各偏好源的专属奖励,并分析统计不确定性如何传播至对偶规划中。将约束目标转化为带KL正则化的拉格朗日函数,其原始最优解为吉布斯策略,从而将学习问题简化为凸对偶问题。提出一种仅优化对偶变量的算法,确保约束以高概率满足,并首次为离线约束偏好学习提供有限样本性能保证。最后,将理论扩展至支持多约束及一般f-散度正则化的情形。

原文摘要 · Abstract (English)

We study offline constrained reinforcement learning from human feedback with multiple preference oracles. Motivated by applications that trade off performance with safety or fairness, we aim to maximize target population utility subject to a minimum protected group welfare constraint. From pairwise comparisons collected under a reference policy, we estimate oracle-specific rewards via maximum likelihood and analyze how statistical uncertainty propagates through the dual program. We cast the constrained objective as a KL-regularized Lagrangian whose primal optimizer is a Gibbs policy, reducing learning to a convex dual problem. We propose a dual-only algorithm that ensures high-probability constraint satisfaction and provide the first finite-sample performance guarantees for offline constrained preference learning. Finally, we extend our theoretical analysis to accommodate multiple constraints and general f-divergence regularization.

强化学习离线学习偏好学习约束优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。