arXiv:2502.20957cs.LG2025-02ICLR被引 1

提出一种新方法,让多目标强化学习能高效处理超多目标场景。

Reward Dimension Reduction for Scalable Multi-Objective Reinforcement Learning

  • 用动态维度压缩技术,在线学习中保留最优解集
  • 在16个目标环境中表现远超现有方法
  • 适合需要扩展多目标能力的研究者

本文提出一种简单有效的奖励维度缩减方法,解决多目标强化学习算法的可扩展性挑战。现有方法大多仅支持2到4个目标,难以扩展至更多目标。本文方法针对在线学习设计,虽基于传统降维思想,但专为动态环境优化,确保变换后仍保持帕累托最优性。我们构建了新的训练与评估框架,实证表明该方法在包含16个目标的环境中显著优于现有在线降维方法,大幅提升学习效率与策略性能。

原文摘要 · Abstract (English)

In this paper, we introduce a simple yet effective reward dimension reduction method to tackle the scalability challenges of multi-objective reinforcement learning algorithms. While most existing approaches focus on optimizing two to four objectives, their abilities to scale to environments with more objectives remain uncertain. Our method uses a dimension reduction approach to enhance learning efficiency and policy performance in multi-objective settings. While most traditional dimension reduction methods are designed for static datasets, our approach is tailored for online learning and preserves Pareto-optimality after transformation. We propose a new training and evaluation framework for reward dimension reduction in multi-objective reinforcement learning and demonstrate the superiority of our method in environments including one with sixteen objectives, significantly outperforming existing online dimension reduction methods.

多目标强化学习维度压缩在线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。