无需种群的高效多策略多目标强化学习方法
Population-Free Pareto Tracking for Sample-Efficient Multi-Policy MORL
- 不依赖进化种群,用极端策略初始化追踪帕累托前沿
- 在6个机器人任务中提升超体积指标,减少交互次数
- 适合追求样本效率的多目标强化学习研究者
多目标强化学习(MORL)是解决现实世界中多冲突目标决策问题的基础框架。现有多策略(MP)方法通常依赖在线进化框架,需维护大规模策略种群,导致高样本复杂度和过多的智能体-环境交互。为缓解此问题,我们提出无种群多策略帕累托前沿追踪(MPFT)框架。该框架不依赖自演化种群,通过单目标极端策略初始化高效的帕累托追踪机制,以追踪帕累托前沿,并进一步对稀疏区域进行稠密化,实现对完整帕累托前沿的精确逼近。MPFT可无缝集成于先进离线MORL算法,显著提升样本效率。我们在最多三个目标的六个机器人控制任务及超过三个目标的三个高维任务上评估了MPFT。实验结果表明,MPFT在超体积和期望效用方面优于现有最优基线,同时显著减少智能体-环境交互次数。这些结果进一步证明,MPFT是一种通用框架,可无缝整合在线与离线MORL算法。
原文摘要 · Abstract (English)
Multi-objective reinforcement learning (MORL) is a fundamental framework for real-world decision-making problems involving multiple conflicting criteria. Existing multi-policy (MP) methods typically rely on online evolutionary frameworks that maintain large policy populations, leading to high sample complexity and excessive agent-environment interactions. To mitigate these limitations, we present Multi-policy Pareto Front Tracking (MPFT), a framework without a self-evolving population. It leverages an efficient Pareto-tracking mechanism initialized with single-objective extreme policies to trace the Pareto front, and further densifies sparse regions to achieve an accurate approximation of the full Pareto front. MPFT can be seamlessly integrated with advanced offline MORL algorithms, thereby substantially improving sample efficiency. We evaluate MPFT on six robotic control tasks with up to three objectives and three high-dimensional tasks with more than three objectives. Experimental results show that MPFT outperforms state-of-the-art baselines in terms of hypervolume and expected utility. It also significantly reduces agent-environment interactions. These results further demonstrate that MPFT serves as a general-purpose framework that can seamlessly integrate both online and offline MORL algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。