用无奖励学习提升多目标强化学习,让智能体更灵活适应用户偏好。
A Reward-Free Viewpoint on Multi-Objective Reinforcement Learning
- 将无奖励强化学习作为辅助任务,提升多目标策略泛化能力。
- 在多个MO-Gymnasium任务上超越现有方法,性能与数据效率双优。
- 适合需要快速适配不同用户偏好的复杂决策场景。
许多序列决策任务需同时优化多个冲突目标,要求策略能根据用户偏好动态调整。多目标强化学习(MORL)主流方法是训练一个以偏好加权奖励为条件的单一策略网络。本文提出新视角:利用无奖励强化学习(RFRL)解决MORL问题。尽管RFRL传统上独立于MORL研究,但其能学习任意奖励函数下的最优策略,天然契合处理未知用户偏好的挑战。我们提出将RFRL的训练目标作为辅助任务,增强MORL,实现超越初始多目标奖励之外的知识共享。为此,我们适配了一种先进的RFRL算法至MORL场景,并引入偏好引导探索策略,聚焦环境中的相关区域。通过大量实验与消融分析,结果表明该方法在多样化的MO-Gymnasium任务中显著优于当前最优的MORL方法,展现出更优性能与数据效率。本工作首次系统性地将RFRL应用于MORL,证明其作为可扩展且实证有效的多目标策略学习方案的潜力。
原文摘要 · Abstract (English)
Many sequential decision-making tasks involve optimizing multiple conflicting objectives, requiring policies that adapt to different user preferences. In multi-objective reinforcement learning (MORL), one widely studied approach} addresses this by training a single policy network conditioned on preference-weighted rewards. In this paper, we explore a novel algorithmic perspective: leveraging reward-free reinforcement learning (RFRL) for MORL. While RFRL has historically been studied independently of MORL, it learns optimal policies for any possible reward function, making it a natural fit for MORL's challenge of handling unknown user preferences. We propose using the RFRL's training objective as an auxiliary task to enhance MORL, enabling more effective knowledge sharing beyond the multi-objective reward function given at training time. To this end, we adapt a state-of-the-art RFRL algorithm to the MORL setting and introduce a preference-guided exploration strategy that focuses learning on relevant parts of the environment. Through extensive experiments and ablation studies, we demonstrate that our approach significantly outperforms the state-of-the-art MORL methods across diverse MO-Gymnasium tasks, achieving superior performance and data efficiency. This work provides the first systematic adaptation of RFRL to MORL, demonstrating its potential as a scalable and empirically effective solution to multi-objective policy learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。