用强化学习动态优化多智能体协调约束,提升导航适应性。
ReCoDe: Reinforcement Learning-based Dynamic Constraint Design for Multi-Agent Coordination
- 基于强化学习动态生成额外约束,增强原有控制器
- 在复杂场景下显著降低拥堵,性能优于手工设计与纯强化学习方法
- 适合需实时调整的多机器人协同任务,如仓储、救援
基于约束的优化是机器人控制的核心,可可靠编码任务与安全需求,如避障或队形保持。然而,在需要复杂协调的多智能体场景中,手工设计的约束可能失效。我们提出 ReCoDe——基于强化学习的动态约束设计,一种去中心化、混合式框架,融合了优化控制的可靠性与多智能体强化学习的适应性。不抛弃专家控制器,而是通过学习额外的动态约束来改进它们,例如在密集场景中约束移动以避免拥堵。通过局部通信,智能体共同限制自身行动空间,在变化条件下更有效协调。本文聚焦于需复杂上下文行为与共识的多智能体导航任务,实验证明 ReCoDe 优于纯手工控制器、其他混合方法及标准 MARL 基线。我们提供了实证(真实机器人)与理论证据:保留用户定义的控制器,即使不完美,也比从零学习更高效,因为 ReCoDe 可动态调节对原控制器的依赖程度。
原文摘要 · Abstract (English)
Constraint-based optimization is a cornerstone of robotics, enabling the design of controllers that reliably encode task and safety requirements such as collision avoidance or formation adherence. However, handcrafted constraints can fail in multi-agent settings that demand complex coordination. We introduce ReCoDe--Reinforcement-based Constraint Design--a decentralized, hybrid framework that merges the reliability of optimization-based controllers with the adaptability of multi-agent reinforcement learning. Rather than discarding expert controllers, ReCoDe improves them by learning additional, dynamic constraints that capture subtler behaviors, for example, by constraining agent movements to prevent congestion in cluttered scenarios. Through local communication, agents collectively constrain their allowed actions to coordinate more effectively under changing conditions. In this work, we focus on applications of ReCoDe to multi-agent navigation tasks requiring intricate, context-based movements and consensus, where we show that it outperforms purely handcrafted controllers, other hybrid approaches, and standard MARL baselines. We give empirical (real robot) and theoretical evidence that retaining a user-defined controller, even when it is imperfect, is more efficient than learning from scratch, especially because ReCoDe can dynamically change the degree to which it relies on this controller.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。