用时间因果知识加速去中心化多智能体强化学习
Decentralizing Multi-Agent Reinforcement Learning with Temporal Causal Information
- 引入时间因果符号知识提升策略兼容性
- 实验证明可显著加快多智能体学习速度
- 适合隐私敏感、通信受限的协作场景
强化学习算法能为单个智能体找到完成特定任务的最优策略,但许多现实问题需要多个智能体协作达成共同目标。例如,仓库中的机器人可能需要无人机协助取高处物品。在去中心化多智能体强化学习(DMARL)中,各智能体独立学习,执行时组合策略,但需满足局部策略兼容性约束,以确保联合后能完成全局任务。本文研究如何向智能体提供高层符号知识,以应对该场景下的独特挑战,如隐私限制、通信瓶颈和性能问题。我们扩展了用于验证局部策略与团队任务兼容性的形式化工具,使具备理论保证的去中心化训练适用于更多场景。此外,实证表明,关于环境事件时间演化的符号知识可显著加速DMARL的学习过程。
原文摘要 · Abstract (English)
Reinforcement learning (RL) algorithms can find an optimal policy for a single agent to accomplish a particular task. However, many real-world problems require multiple agents to collaborate in order to achieve a common goal. For example, a robot executing a task in a warehouse may require the assistance of a drone to retrieve items from high shelves. In Decentralized Multi-Agent RL (DMARL), agents learn independently and then combine their policies at execution time, but often must satisfy constraints on compatibility of local policies to ensure that they can achieve the global task when combined. In this paper, we study how providing high-level symbolic knowledge to agents can help address unique challenges of this setting, such as privacy constraints, communication limitations, and performance concerns. In particular, we extend the formal tools used to check the compatibility of local policies with the team task, making decentralized training with theoretical guarantees usable in more scenarios. Furthermore, we empirically demonstrate that symbolic knowledge about the temporal evolution of events in the environment can significantly expedite the learning process in DMARL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。