不用复杂奖励,自监督目标达成让多智能体自动协作探索。
Self-Supervised Goal-Reaching Results in Multi-Agent Cooperation and Exploration
- 用自监督目标达成替代传统奖励函数,引导智能体学习
- 在稀疏奖励下性能超越已有方法,部分任务成功率超50%
- 无需显式探索机制,能发现复杂协作策略,适合复杂环境
为使自主智能体群体达成特定目标,需进行协调与长时程推理。本文探讨在多智能体设置中实现有效协调与探索所需的最小条件,提出自监督目标达成方法:智能体通过最大化访问目标状态的概率来学习,而非依赖复杂奖励函数或显式协作机制。尽管反馈信号稀疏,实验表明该方法仍能有效学习。在多智能体强化学习基准测试中,该方法优于其他仅拥有相同稀疏奖励信号的方案。进一步实验证明,多智能体自监督目标达成比单智能体策略更具鲁棒性。即使无显式探索机制,该方法在其他方法均无法成功的情况下,仍能探索出非平凡的中间协作策略。
原文摘要 · Abstract (English)
For groups of autonomous agents to achieve a particular goal, they must engage in coordination and long-horizon reasoning. Rather than relying on complex reward functions and explicit cooperation mechanisms, we ask what minimal ingredients are required for effective coordination and exploration to emerge in multi-agent settings. We investigate this question through self-supervised goal-reaching, where agents aim to maximize the likelihood of visiting a goal state rather than maximizing a reward. Despite a sparse feedback signal, we present empirical results that show self-supervised goal-reaching techniques enable agents to learn from such feedback. On MARL benchmarks, self-supervised goal-reaching outperforms alternative approaches that have access to the same sparse reward signal. Furthermore, we empirically demonstrate that multi-agent self-supervised goal-reaching approaches can be more robust than single-agent strategies. While there is no explicit exploration mechanism, this approach explores nontrivial intermediate coordination strategies in sparse settings where alternative approaches fail to achieve a single success.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。