提出新方法提升多目标强化学习的最优策略逼近效果
PA2D-MORL: Pareto Ascent Directional Decomposition based Multi-Objective Reinforcement Learning
- 基于帕累托上升方向分解目标,动态调整权重优化策略
- 在多个机器人控制任务中显著提升解集质量与稳定性
- 适合需要平衡多目标的复杂决策场景
多目标强化学习(MORL)为处理冲突目标的决策问题提供了有效方案,但在连续或高维状态-动作空间的复杂任务中,仍难以获得高质量的帕累托策略集近似。本文提出基于帕累托上升方向分解的多目标强化学习方法(PA2D-MORL),通过帕累托上升方向选择标量化权重,并计算多目标策略梯度,确定联合优化方向,确保所有目标同时提升。同时,在进化框架下对多个策略进行选择性优化,从不同方向逼近帕累托前沿。此外,引入帕累托自适应微调策略,增强帕累托前沿近似的密度与分布范围。在多个多目标机器人控制任务上的实验表明,该方法在结果质量和稳定性方面均显著优于当前最先进算法。
原文摘要 · Abstract (English)
Multi-objective reinforcement learning (MORL) provides an effective solution for decision-making problems involving conflicting objectives. However, achieving high-quality approximations to the Pareto policy set remains challenging, especially in complex tasks with continuous or high-dimensional state-action space. In this paper, we propose the Pareto Ascent Directional Decomposition based Multi-Objective Reinforcement Learning (PA2D-MORL) method, which constructs an efficient scheme for multi-objective problem decomposition and policy improvement, leading to a superior approximation of Pareto policy set. The proposed method leverages Pareto ascent direction to select the scalarization weights and computes the multi-objective policy gradient, which determines the policy optimization direction and ensures joint improvement on all objectives. Meanwhile, multiple policies are selectively optimized under an evolutionary framework to approximate the Pareto frontier from different directions. Additionally, a Pareto adaptive fine-tuning approach is applied to enhance the density and spread of the Pareto frontier approximation. Experiments on various multi-objective robot control tasks show that the proposed method clearly outperforms the current state-of-the-art algorithm in terms of both quality and stability of the outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。