arXiv:2501.06773cs.LG2025-01AAAI被引 38

用超网络生成多目标强化学习策略,实现高效精准的帕累托前沿覆盖。

Pareto Set Learning for Multi-Objective Reinforcement Learning

  • 通过超网络为不同权重生成专属策略参数,实现多目标分解求解。
  • 在多个基准上显著提升超体积与稀疏性指标,覆盖更密集的帕累托前沿。
  • 适用于任何强化学习算法,适合需个性化决策的机器人、游戏等场景。

多目标决策问题广泛存在于视频游戏、导航和机器人等领域。尽管强化学习在优化决策方面具有优势,但现有多目标强化学习方法要么无法获得完整的帕累托前沿,要么仅使用单一策略网络应对所有目标偏好,难以生成个性化解。为此,本文提出基于分解的新型框架Pareto Set Learning for MORL(PSL-MORL),利用超网络为每个分解权重生成策略网络参数,高效生成针对不同加权子问题的差异化策略。该框架通用性强,兼容任意强化学习算法。理论分析证明其模型容量更优且策略最优。大量实验表明,PSL-MORL在多个基准上实现了帕累托前沿的密集覆盖,显著优于当前最先进的MORL方法,在超体积与稀疏性指标上表现突出。

原文摘要 · Abstract (English)

Multi-objective decision-making problems have emerged in numerous real-world scenarios, such as video games, navigation and robotics. Considering the clear advantages of Reinforcement Learning (RL) in optimizing decision-making processes, researchers have delved into the development of Multi-Objective RL (MORL) methods for solving multi-objective decision problems. However, previous methods either cannot obtain the entire Pareto front, or employ only a single policy network for all the preferences over multiple objectives, which may not produce personalized solutions for each preference. To address these limitations, we propose a novel decomposition-based framework for MORL, Pareto Set Learning for MORL (PSL-MORL), that harnesses the generation capability of hypernetwork to produce the parameters of the policy network for each decomposition weight, generating relatively distinct policies for various scalarized subproblems with high efficiency. PSL-MORL is a general framework, which is compatible for any RL algorithm. The theoretical result guarantees the superiority of the model capacity of PSL-MORL and the optimality of the obtained policy network. Through extensive experiments on diverse benchmarks, we demonstrate the effectiveness of PSL-MORL in achieving dense coverage of the Pareto front, significantly outperforming state-of-the-art MORL methods in the hypervolume and sparsity indicators.

多目标强化学习帕累托前沿超网络策略生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。