arXiv:2412.20361cs.MAcs.AI2024-12被引 4

用熵激励探索,让多智能体在安全约束下更高效协作。

Safe Multiagent Coordination via Entropic Exploration

  • 通过最大化观测熵鼓励探索,突破安全约束下的行为僵局。
  • 在复杂场景中任务表现优于或媲美基线,不安全行为减少50%。
  • 适合需要安全协作的多智能体系统,如自动驾驶、机器人群组。

许多现实世界的多智能体学习问题涉及安全约束。传统安全强化学习算法会限制智能体行为,从而抑制探索——发现有效协同策略的关键环节。此外,现有研究通常为每个智能体单独设置约束,尚未探讨团队联合约束的优势。本文从理论与实践角度分析团队约束,并提出用于约束型多智能体强化学习的熵激励探索方法(E2C),通过最大化观测熵来激励探索,促进安全且高效的协同行为学习。在日益复杂的多个领域实验中,E2C智能体在任务表现上达到或超越常见无约束及有约束基线,同时将不安全行为降低最多达50%。

原文摘要 · Abstract (English)

Many real-world multiagent learning problems involve safety concerns. In these setups, typical safe reinforcement learning algorithms constrain agents' behavior, limiting exploration -- a crucial component for discovering effective cooperative multiagent behaviors. Moreover, the multiagent literature typically models individual constraints for each agent and has yet to investigate the benefits of using joint team constraints. In this work, we analyze these team constraints from a theoretical and practical perspective and propose entropic exploration for constrained multiagent reinforcement learning (E2C) to address the exploration issue. E2C leverages observation entropy maximization to incentivize exploration and facilitate learning safe and effective cooperative behaviors. Experiments across increasingly complex domains show that E2C agents match or surpass common unconstrained and constrained baselines in task performance while reducing unsafe behaviors by up to $50\%$.

多智能体安全强化学习熵激励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。