arXiv:2502.18699cs.CLcs.LG2025-02ICML被引 10

用加权融合方式高效整合多种偏好,提升大模型对多元人类偏好的适应性。

MPO: An Efficient Post-Processing Framework for Mixing Diverse Preference Alignment

  • 通过批量随机镜面下降法计算策略权重,对已有策略进行对数线性融合。
  • 在多偏好场景下表现均衡,性能优于或相当主流方法,计算成本显著降低。
  • 适合追求高效、稳定对齐且不想从头训练的开发者使用。

基于人类反馈的强化学习(RLHF)在对齐大语言模型(LLMs)方面展现出潜力,但其依赖单一奖励模型常忽视人类偏好的多样性。近期方法通过多维反馈微调对应奖励模型并使用强化学习训练LLMs,但过程成本高且不稳定,尤其在人类偏好存在竞争与异质性时。本文提出混合偏好优化(MPO),一种后处理框架,用于聚合单目标策略,作为多目标RLHF(MORLHF)和最大最小RLHF(MaxMin-RLHF)的替代方案。MPO避免从零开始对齐,而是通过批量随机镜面下降计算各策略权重,对现有策略进行对数线性融合,生成统一策略。实证结果表明,MPO在多种偏好间实现平衡表现,性能优于或匹配现有模型,且计算成本大幅降低。

原文摘要 · Abstract (English)

Reinforcement Learning from Human Feedback (RLHF) has shown promise in aligning large language models (LLMs). Yet its reliance on a singular reward model often overlooks the diversity of human preferences. Recent approaches address this limitation by leveraging multi-dimensional feedback to fine-tune corresponding reward models and train LLMs using reinforcement learning. However, the process is costly and unstable, especially given the competing and heterogeneous nature of human preferences. In this paper, we propose Mixing Preference Optimization (MPO), a post-processing framework for aggregating single-objective policies as an alternative to both multi-objective RLHF (MORLHF) and MaxMin-RLHF. MPO avoids alignment from scratch. Instead, it log-linearly combines existing policies into a unified one with the weight of each policy computed via a batch stochastic mirror descent. Empirical results demonstrate that MPO achieves balanced performance across diverse preferences, outperforming or matching existing models with significantly reduced computational costs.

偏好对齐强化学习大模型高效优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。