arXiv:2411.12433cs.AI2024-11被引 2

用偏好引导的梯度变异,高效生成多目标最优解集

Preference-Conditioned Gradient Variations for Multi-Objective Quality-Diversity

  • 基于偏好条件的策略梯度变异,精准探索目标空间
  • 在6个机器人任务中优于或持平现有最优方法
  • 生成更平滑的权衡解,适合复杂多目标优化场景

在机器人、金融等多个领域,质量-多样性算法被用于生成多样且高性能的解集。多目标质量-多样性算法成为解决复杂多目标问题的有力工具。然而,现有方法受限于搜索能力:例如多目标地图精英法依赖随机遗传变异,在高维空间表现不佳;尽管已有基于梯度的变异提升效率,但通常仅单独优化各目标,无法实现理想的权衡。本文提出偏好条件策略梯度与拥挤机制的多目标地图精英算法(Preference-Conditioned Policy-Gradient and Crowding Mechanisms for Multi-Objective Map-Elites),利用偏好引导的策略梯度变异高效发现目标空间中的优质区域,并通过拥挤机制确保解在非支配前沿上均匀分布。我们在六个机器人运动任务上评估该方法,结果表明其在所有任务中均优于或匹配当前最先进的多目标质量-多样性方法,包括两个新提出的三目标任务。更重要的是,基于新提出的稀疏性度量,本方法生成的权衡解更加平滑。

原文摘要 · Abstract (English)

In a variety of domains, from robotics to finance, Quality-Diversity algorithms have been used to generate collections of both diverse and high-performing solutions. Multi-Objective Quality-Diversity algorithms have emerged as a promising approach for applying these methods to complex, multi-objective problems. However, existing methods are limited by their search capabilities. For example, Multi-Objective Map-Elites depends on random genetic variations which struggle in high-dimensional search spaces. Despite efforts to enhance search efficiency with gradient-based mutation operators, existing approaches consider updating solutions to improve on each objective separately rather than achieving desired trade-offs. In this work, we address this limitation by introducing Multi-Objective Map-Elites with Preference-Conditioned Policy-Gradient and Crowding Mechanisms: a new Multi-Objective Quality-Diversity algorithm that uses preference-conditioned policy-gradient mutations to efficiently discover promising regions of the objective space and crowding mechanisms to promote a uniform distribution of solutions on the non-dominated front. We evaluate our approach on six robotics locomotion tasks and show that our method outperforms or matches all state-of-the-art Multi-Objective Quality-Diversity methods in all six, including two newly proposed tri-objective tasks. Importantly, our method also achieves a smoother set of trade-offs, as measured by newly-proposed sparsity-based metrics.

多目标优化质量多样性强化学习机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。