arXiv:2506.05619cs.AIcs.LG2025-06被引 4

让AI更公平地学习大众意见,避免少数声音主导决策。

Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic Framework

  • 基于社会选择理论,从对比数据推断评价者偏好分布
  • 确保策略符合比例对齐与抗操纵等核心公理要求
  • 适用于推荐系统和大模型对齐,兼顾公平性与可扩展性

传统偏好学习方法在聚合多评估者意见时倾向于优先考虑更普遍的观点,可能导致政策偏向特定观点或群体,并易受策略性操纵。为此,我们提出一种新型偏好学习框架,能够使聚合意见与策略的比例与真实评价者群体偏好分布相一致。该方法基于社会选择理论,直接从成对比较数据中推断评价者群体分布的可行集。利用这些估计,算法构建满足社会选择基本公理(单调性、帕累托效率)及新提出的群体比例对齐与群体有界可操纵性公理的策略。此外,我们提出一种软最大松弛方法,可平滑权衡群体比例对齐与选择胜过所有其他选项的康多塞赢家(Condorcet winner)。最后,我们在表格推荐任务和大语言模型对齐上验证了该方法的有效性与可扩展性。

原文摘要 · Abstract (English)

Conventional preference learning methods often prioritize opinions held more widely when aggregating preferences from multiple evaluators. This may result in policies that are biased in favor of some types of opinions or groups and susceptible to strategic manipulation. To address this issue, we develop a novel preference learning framework capable of aligning aggregate opinions and policies proportionally with the true population distribution of evaluator preferences. Grounded in social choice theory, our approach infers the feasible set of evaluator population distributions directly from pairwise comparison data. Using these estimates, the algorithm constructs a policy that satisfies foundational axioms from social choice theory, namely monotonicity and Pareto efficiency, as well as our newly-introduced axioms of population-proportional alignment and population-bounded manipulability. Moreover, we propose a soft-max relaxation method that smoothly trades off population-proportional alignment with the selection of the Condorcet winner (which beats all other options in pairwise comparisons). Finally, we validate the effectiveness and scalability of our approach through experiments on both tabular recommendation tasks and large language model alignment.

偏好学习公平对齐社会选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。