在联邦学习中用自适应策略更公平地对齐大模型与多元人类偏好。
A Systematic Evaluation of Preference Aggregation in Federated RLHF for Pluralistic Alignment of LLMs
- 通过动态调整各群体偏好权重,实现更公平的奖励聚合。
- 实验表明新方法在保持良好对齐效果的同时显著提升公平性。
- 适合关注多群体公平对齐的AI伦理与联邦学习研究者。
本文针对联邦学习环境中大语言模型(LLMs)对多元人类偏好的对齐挑战,提出一个系统评估框架,用于分析不同偏好聚合策略在对齐质量与公平性之间的权衡。在联邦设置中,各群体本地评估生成结果并产生奖励信号,服务器仅聚合群体级奖励而不访问原始数据。我们评估了标准聚合方法(最小值、最大值、平均值),并引入一种新型自适应方案,根据群体历史对齐表现动态调整偏好权重。基于PPO的强化学习人类反馈(RLHF)流程在问答任务上的实验表明,该自适应方法在保持竞争力对齐得分的同时,持续实现更优的公平性。本工作为跨多元人群评估大模型行为提供了稳健方法,并为开发真正多元且公平对齐的模型提供实用解决方案。
原文摘要 · Abstract (English)
This paper addresses the challenge of aligning large language models (LLMs) with diverse human preferences within federated learning (FL) environments, where standard methods often fail to adequately represent diverse viewpoints. We introduce a comprehensive evaluation framework that systematically assesses the trade-off between alignment quality and fairness when using different aggregation strategies for human preferences. In our federated setting, each group locally evaluates rollouts and produces reward signals, and the server aggregates these group-level rewards without accessing any raw data. Specifically, we evaluate standard reward aggregation techniques (min, max, and average) and introduce a novel adaptive scheme that dynamically adjusts preference weights based on a group's historical alignment performance. Our experiments on question-answering (Q/A) tasks using a PPO-based RLHF pipeline demonstrate that our adaptive approach consistently achieves superior fairness while maintaining competitive alignment scores. This work offers a robust methodology for evaluating LLM behavior across diverse populations and provides a practical solution for developing truly pluralistic and fairly aligned models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。