让多机器人在竞争目标中保持协调,支持去中心化部署。
Sampling-Based Coordination-Informed Multi-Objective Multi-Robot Reinforcement Learning

- 用采样方法融合全局信息训练,实现去中心化决策。
- 相比顶尖方法,超体积提升21.2%,策略更稳定。
- 适合资源分配与攻防对抗等真实场景,抗部分观测噪声。
多机器人系统需在竞争目标间权衡并保持协调行为。现有方法常依赖固定或集中式协调,限制适应性且违反分布式约束。本文提出协调感知多目标强化学习框架(CIMORL),融合分布式权重预测、特权专家训练策略及帕累托最优的理论保证。提出基础CIMORL方法及其两种基于采样的变体:CIMORL-TS(树搜索)和CIMORL-MPPI(MPPI),利用训练阶段的全局信息实现完全去中心化部署。在协作与对抗场景中的实验表明,相较最先进基线,超体积提升21.2%,策略稳定性更优。在Crazyflie无人机上的真实世界实验进一步验证了该框架在资源分配与多攻击者-多防御场景下,于部分可观测条件下的鲁棒性。
原文摘要 · Abstract (English)
Multi-robot systems must simultaneously optimize competing objectives while maintaining coordinated behavior. Existing multi-agent reinforcement learning approaches often rely on fixed or centralized coordination, which limits adaptability and violates distributed constraints. This work introduces the Coordination-Informed Multi-Objective Reinforcement Learning (CIMORL) framework, integrating a distributed weight prediction mechanism, a privileged expert training strategy, and theoretical guarantees for Pareto-optimal solutions. We present the base CIMORL method alongside two sampling-based variants, CIMORL-TS (Tree Search) and CIMORL-MPPI (MPPI), which leverage privileged global information during training to enable fully decentralized deployment. Experimental validation in cooperative and adversarial scenarios demonstrates a $21.2\%$ hypervolume improvement and superior policy stability compared to state-of-the-art baselines. Real-world experiments with Crazyflie drones further validate the framework's robustness in resource allocation and multi-attacker multi-defend scenarios under partial observability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。