arXiv:2507.14066cs.LG2025-07中稿 · publication in IEE…被引 15

用用户偏好替代复杂奖励函数,实现多目标强化学习的高效优化

Preference-based Multi-Objective Reinforcement Learning

  • 通过用户偏好构建多目标奖励模型,避免人工设计奖励函数
  • 理论证明该方法能覆盖全部帕累托最优解,实验表现优于已知最优奖励方法
  • 适合需要灵活权衡多个冲突目标的实际场景,如自动驾驶与能源管理

多目标强化学习(MORL)是一种用于优化多目标任务的结构化方法。然而,传统方法通常依赖预先定义的奖励函数,难以平衡冲突目标,且易导致简化。偏好作为更灵活直观的决策指导,可替代复杂奖励设计。本文提出基于偏好的多目标强化学习(Pb-MORL),将偏好形式化融入MORL框架。理论上证明偏好可导出整个帕累托前沿的策略。为此,我们构建与给定偏好一致的多目标奖励模型,并提供理论证明:优化该模型等价于训练帕累托最优策略。在基准多目标任务、多能源管理任务及多车道高速公路自动驾驶任务上的大量实验表明,本方法表现优异,超越使用真实奖励函数的基线方法,展现出在复杂现实系统中的应用潜力。

原文摘要 · Abstract (English)

Multi-objective reinforcement learning (MORL) is a structured approach for optimizing tasks with multiple objectives. However, it often relies on pre-defined reward functions, which can be hard to design for balancing conflicting goals and may lead to oversimplification. Preferences can serve as more flexible and intuitive decision-making guidance, eliminating the need for complicated reward design. This paper introduces preference-based MORL (Pb-MORL), which formalizes the integration of preferences into the MORL framework. We theoretically prove that preferences can derive policies across the entire Pareto frontier. To guide policy optimization using preferences, our method constructs a multi-objective reward model that aligns with the given preferences. We further provide theoretical proof to show that optimizing this reward model is equivalent to training the Pareto optimal policy. Extensive experiments in benchmark multi-objective tasks, a multi-energy management task, and an autonomous driving task on a multi-line highway show that our method performs competitively, surpassing the oracle method, which uses the ground truth reward function. This highlights its potential for practical applications in complex real-world systems.

强化学习多目标优化偏好学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。