arXiv:2607.29559cs.AIcs.RO2026-07

让智能体通过人类偏好学习多目标平衡策略

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback

论文配图:LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback
图 1 · 摘自论文原文
  • 从人类偏好中联合学习多个目标的奖励模型与策略
  • 在多任务基准上优于现有基线方法
  • 适合无预设奖励函数的复杂决策场景

强化学习系统通常依赖单一明确的标量奖励函数进行训练。然而,现实决策任务常涉及性能与效率等多重竞争目标,真实奖励函数难以定义或不可获取。虽然多目标强化学习(MORL)通过向量形式建模奖励以处理此类权衡,但现有方法仍需每个目标的明确奖励函数,继承了单目标强化学习的局限性。与此同时,基于偏好的强化学习(PbRL)在无需预定义奖励函数的情况下,通过人类反馈学习奖励,展现出解决复杂任务的巨大潜力,但主要局限于单目标设置。本文提出LEMUR:一种结合人类偏好与多目标强化学习的新框架,使智能体通过与多位人类交互,从偏好反馈中学习最优的多目标策略。该方法联合学习策略和多个目标特定的奖励模型,实现学习过程中对竞争目标的有效平衡。我们在多种基准多目标任务上评估LEMUR,实证结果表明其性能显著优于基线方法。本方法为在无预定义奖励函数的情况下解决多目标决策问题提供了有前景的方向。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) systems are typically trained using a single, well-specified scalar reward function. However, real-world decision-making tasks often involve multiple, competing objectives, such as performance versus efficiency, where ground-truth reward functions are difficult to specify or inaccessible. While Multi-Objective RL (MORL) addresses such trade-offs by modeling rewards as vectors, existing approaches typically assume access to a well-specified reward function for each objective, inheriting the same challenges faced by single-objective RL. Meanwhile, Preference-based RL (PbRL) has shown great potential in solving complex tasks without access to a pre-defined reward function through reward learning from human feedback, yet has largely been studied in single-objective settings. In this work, we bridge this gap with LEMUR: Learning to Align with Multi-Objective Reinforcement Learning with Preference feedback, a novel framework where an agent interactively learns from the preferences of multiple humans to learn optimal multi-objective policies. Our approach jointly learns policies and multiple objective-specific reward models from human feedback, enabling agents to effectively balance competing objectives during learning. We evaluate LEMUR on a variety of benchmark multi-objective tasks, and empirical results demonstrate its superior performance over baseline methods. Our method presents a promising direction for solving multi-objective decision-making tasks without pre-defined reward functions.

多目标强化学习偏好学习人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。