用关系规划器提升多智能体强化学习的样本效率与泛化能力
Combining Planning and Reinforcement Learning for Solving Relational Multiagent Domains
- 将关系规划器作为中心控制器,结合状态抽象
- 在复杂关系场景中实现高效采样与任务迁移
- 适合需要通用策略的多智能体系统设计
多智能体强化学习(MARL)因状态和动作空间呈指数增长,以及环境非平稳性,面临显著挑战,导致样本效率低下,并阻碍跨任务泛化。这一问题在关系型设置中尤为突出,领域知识虽关键但常被现有MARL算法忽视。为此,我们提出将关系规划器作为中心控制器,结合高效的态抽象与强化学习。该方法展现出良好的样本效率,有效促进任务迁移与泛化。
原文摘要 · Abstract (English)
Multiagent Reinforcement Learning (MARL) poses significant challenges due to the exponential growth of state and action spaces and the non-stationary nature of multiagent environments. This results in notable sample inefficiency and hinders generalization across diverse tasks. The complexity is further pronounced in relational settings, where domain knowledge is crucial but often underutilized by existing MARL algorithms. To overcome these hurdles, we propose integrating relational planners as centralized controllers with efficient state abstractions and reinforcement learning. This approach proves to be sample-efficient and facilitates effective task transfer and generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。