让大模型更好适应不同用户偏好,提升个性化对齐效果。
Personalized Group Relative Policy Optimization for Heterogenous Preference Alignment
- 基于用户偏好分组,用历史奖励数据归一化优势值,避免群体偏见。
- 在多个任务上收敛更快、奖励更高,显著提升少数偏好学习能力。
- 适合需要精准个性化对齐的应用场景,如定制助手或内容推荐。
尽管大型语言模型具备强大的通用能力,但标准后训练方法(如基于人类反馈的强化学习)通常优化单一全局目标,难以适配多样化的个体偏好。现有组相对策略优化(GRPO)框架依赖于样本可交换性假设,导致不同用户奖励分布混淆,系统性偏向主流偏好而压制少数信号。为此,我们提出个性化GRPO(P-GRPO),通过将优势值归一化基于偏好组的历史奖励而非当前批次统计,解耦优势估计与即时批量数据。该设计保留了学习差异偏好的对比信号。实验表明,P-GRPO在多种任务中均实现更快收敛和更高奖励,显著增强对异质偏好信号的恢复与对齐能力。结果证明,在优化层面考虑奖励异质性,是实现与多样化人类偏好一致且不牺牲通用能力的关键。
原文摘要 · Abstract (English)
Despite their sophisticated general-purpose capabilities, Large Language Models (LLMs) often fail to align with diverse individual preferences because standard post-training methods, like Reinforcement Learning with Human Feedback (RLHF), optimize for a single, global objective. While Group Relative Policy Optimization (GRPO) is a widely adopted on-policy reinforcement learning framework, its group-based normalization implicitly assumes that all samples are exchangeable, inheriting this limitation in personalized settings. This assumption conflates distinct user reward distributions and systematically biases learning toward dominant preferences while suppressing minority signals. To address this, we introduce Personalized GRPO (P-GRPO), a novel alignment framework that decouples advantage estimation from immediate batch statistics. By normalizing advantages against preference-group-specific reward histories rather than the concurrent generation group, P-GRPO preserves the contrastive signal necessary for learning distinct preferences. We evaluate P-GRPO across diverse tasks and find that it consistently achieves faster convergence and higher rewards than standard GRPO, thereby enhancing its ability to recover and align with heterogeneous preference signals. Our results demonstrate that accounting for reward heterogeneity at the optimization level is essential for building models that faithfully align with diverse human preferences without sacrificing general capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。