arXiv:2503.00295cs.CLcs.LG2025-03AAAI被引 24

让大模型同时满足多种用户偏好,还能实时调整。

Robust Multi-Objective Preference Alignment with Online DPO

  • 用在线策略优化实现多目标偏好对齐,训练一次可灵活适应不同偏好组合。
  • 在两个基准测试中均超越现有方法,且推理时可自由调节偏好权重。
  • 适合需要个性化、可配置的AI系统开发者使用。

大语言模型的多目标偏好对齐对于构建可配置、个性化、有用且安全的AI系统至关重要。然而,在推理时以可变权重优化模型输出以满足多样化目标仍具挑战性。现有方法或训练成本过高,或难以有效引导模型行为。本文提出多目标在线DPO(MO-ODPO)算法,能稳健高效地对齐多个可能冲突的人类偏好。该方法引入提示条件机制,仅需训练单一条件化策略,即可在推理时适配新的偏好组合。在两个主流基准上的实验表明,MO-ODPO在帕累托意义上优于现有基线,并展现出优异的推理时可调控性。

原文摘要 · Abstract (English)

Multi-objective preference alignment of large language models (LLMs) is critical for developing AI systems that are more configurable, personalizable, helpful, and safe. However, optimizing model outputs to satisfy diverse objectives with variable weights at inference time for truly personalized models presents a significant challenge. Existing approaches are either computationally expensive to train or do not sufficiently steer model behaviors. This paper introduces the Multi-Objective Online DPO (MO-ODPO) algorithm, designed to robustly and efficiently align model behaviors with multiple, potentially conflicting human preferences. Our approach incorporates a prompt conditioning mechanism, allowing us to train a single preference-conditional policy, that can adapt to new preference combinations at inference. Experiments on two popular benchmarks show that MO-ODPO Pareto-dominates existing baselines while providing excellent inference-time steerability between diverse objectives.

多目标对齐在线学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。