提出首个连续多目标多智能体强化学习的内层框架,实现高效协同优化。
MOMA-AC: A preference-driven actor-critic framework for continuous multi-objective multi-agent reinforcement learning
- 采用多头策略网络与偏好条件批评家,统一建模多智能体帕累托最优解。
- 在协作运动任务中,期望效用和超体积指标显著优于基线方法。
- 适用于需协同处理冲突目标的连续空间多智能体系统,如机器人集群。
本文针对多目标多智能体强化学习(MOMARL)中的关键空白,提出首个专用于连续状态与动作空间的内层演员-评论家框架:多目标多智能体演员-评论家(MOMA-AC)。基于单目标单智能体算法,我们以孪生延迟深度确定性策略梯度(TD3)和深度确定性策略梯度(DDPG)实例化该框架,得到MOMA-TD3与MOMA-DDPG。该框架结合多头演员网络、集中式评论家及目标偏好条件架构,使单一神经网络可编码所有智能体在冲突目标下的帕累托前沿最优策略。我们还通过整合现有的多智能体单目标物理模拟器与其多目标单智能体版本,构建了一个自然的连续MOMARL测试套件。在协作运动任务上的评估显示,该框架在预期效用和超体积指标上均显著优于外层训练和独立训练基线,且随智能体数量增加仍保持稳定可扩展性。结果确立了该框架在连续多智能体多目标策略学习中的基础地位。
原文摘要 · Abstract (English)
This paper addresses a critical gap in Multi-Objective Multi-Agent Reinforcement Learning (MOMARL) by introducing the first dedicated inner-loop actor-critic framework for continuous state and action spaces: Multi-Objective Multi-Agent Actor-Critic (MOMA-AC). Building on single-objective, single-agent algorithms, we instantiate this framework with Twin Delayed Deep Deterministic Policy Gradient (TD3) and Deep Deterministic Policy Gradient (DDPG), yielding MOMA-TD3 and MOMA-DDPG. The framework combines a multi-headed actor network, a centralised critic, and an objective preference-conditioning architecture, enabling a single neural network to encode the Pareto front of optimal trade-off policies for all agents across conflicting objectives in a continuous MOMARL setting. We also outline a natural test suite for continuous MOMARL by combining a pre-existing multi-agent single-objective physics simulator with its multi-objective single-agent counterpart. Evaluating cooperative locomotion tasks in this suite, we show that our framework achieves statistically significant improvements in expected utility and hypervolume relative to outer-loop and independent training baselines, while demonstrating stable scalability as the number of agents increases. These results establish our framework as a foundational step towards robust, scalable multi-objective policy learning in continuous multi-agent domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。