用抽签机制选出代表性人群,让AI更公平地学习公众偏好。
Democratic Preference Alignment via Sortition-Weighted RLHF
- 通过抽签方式选取有代表性的评价者,替代传统方便抽样。
- 小样本抽签组表现最优,大模型上效果提升更明显。
- 适合关注AI伦理与社会公平的研究者和开发者。
AI应学习谁的价值观?基于人类偏好的对齐方法(如RLHF)依赖于人类评分者,但这些评分者通常为便利样本,存在人口统计偏差。本文提出民主偏好优化(DemPO),引入公民大会常用的抽签机制,用于偏好对齐微调。DemPO提供两种训练方案:硬面板仅使用抽签生成的代表性小群体数据;软面板保留全部数据,但按抽签中被选中的概率重新加权每位评分者。理论证明,软面板加权可闭式恢复硬面板目标。在包含评分者人口统计信息及75条独立获取的美国代表性民意宪法的公开偏好数据集上,对10亿至80亿参数的Llama模型进行评估。六种聚合方法下,硬面板始终排名第一,软面板持续优于未加权基线,且效果随模型规模增大而增强。结果表明,在偏好收集阶段即确保人口代表性,比事后修正更能使模型行为反映代表性群体的价值取向。
原文摘要 · Abstract (English)
Whose values should AI systems learn? Preference based alignment methods like RLHF derive their training signal from human raters, yet these rater pools are typically convenience samples that systematically over represent some demographics and under represent others. We introduce Democratic Preference Optimization, or DemPO, a framework that applies algorithmic sortition, the same mechanism used to construct citizen assemblies, to preference based fine tuning. DemPO offers two training schemes. Hard Panel trains exclusively on preferences from a quota satisfying mini public sampled via sortition. Soft Panel retains all data but reweights each rater by their inclusion probability under the sortition lottery. We prove that Soft Panel weighting recovers the expected Hard Panel objective in closed form. Using a public preference dataset that pairs human judgments with rater demographics and a seventy five clause constitution independently elicited from a representative United States panel, we evaluate Llama models from one billion to eight billion parameters fine tuned under each scheme. Across six aggregation methods, the Hard Panel consistently ranks first and the Soft Panel consistently outperforms the unweighted baseline, with effect sizes growing as model capacity increases. These results demonstrate that enforcing demographic representativeness at the preference collection stage, rather than post hoc correction, yields models whose behavior better reflects values elicited from representative publics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。