arXiv:2503.01233cs.CL2025-03

让大模型同时更贴心又不惹祸,还能灵活适应不同用户偏好。

PEO: Improving Bi-Factorial Preference Alignment with Post-Training Policy Extrapolation

  • 通过三阶段流程一次性生成多组最优策略,避免反复训练。
  • 在多个大模型上实验,性能优于基线方法,提升灵活性与效率。
  • 适合需要个性化对齐的场景,如客服、内容创作等应用。

大语言模型与人类价值观对齐面临关键挑战,尤其在兼顾有用性与无害性等矛盾目标时。现有方法如基于人类反馈的强化学习(RLHF)在多目标优化中存在不稳定和低效问题,而直接偏好优化(DPO)缺乏动态权衡机制。为此,我们提出后训练外推优化(PEO),一种新型高效双因素对齐框架。PEO通过三阶段流程——(1)特定方面学习,(2)通过插值初始化通用模型,(3)后训练外推优化——在一次训练中生成一组帕累托最优策略。该方法可在推理时动态适应不同用户偏好,无需重新训练。在多个大模型上的全面实验表明,PEO相比基线方法实现了更优的帕累托前沿,具备更高的灵活性与计算效率。理论分析进一步揭示了PEO克服优化瓶颈的能力,为可扩展、个性化的对齐提供新路径。

原文摘要 · Abstract (English)

The alignment of large language models with human values presents a critical challenge, particularly when balancing conflicting objectives like helpfulness and harmlessness. Existing approaches, such as Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO), face notable limitations: RLHF suffers from instability and inefficiency in multi-objective optimization, while DPO lacks mechanisms for dynamic trade-offs. To address these challenges, we propose Post-Training Extrapolation Optimization (PEO), a novel and efficient framework for bi-factorial alignment. PEO generates a family of Pareto-optimal policies in a single training pass by leveraging a three-phase pipeline: (1) aspect-specific learning, (2) generalist initialization via interpolation, and (3) post-training optimization via extrapolation. PEO enables dynamic adaptation to diverse user preferences at inference time without retraining. Our comprehensive experiments across multiple LLMs demonstrate that PEO achieves superior Pareto fronts compared to baselines, offering improved flexibility and computational efficiency. Theoretical analyses further highlight PEO's capacity to overcome optimization bottlenecks, paving the way for scalable, personalized alignment.

模型对齐偏好优化大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。