无需队友状态,靠观察就能识别团队任务并协作。
RecBayes: Recurrent Bayesian Ad Hoc Teamwork in Large Partially Observable Domains
- 用循环贝叶斯分类器从观测中识别团队与任务
- 在100万状态、2^125种观测的大环境中仍有效
- 适合大规模部分可观测多智能体协作场景
本文提出RecBayes,一种在部分可观测环境下进行即兴团队协作的新方法。该方法无需任何时刻访问环境状态或队友动作,仅通过过往经验训练的循环贝叶斯分类器,即可从观测中有效识别已知团队和正在执行的任务。相比PO-GPL、FEAT等需完全可观测状态或动作的方法,以及需表格化建模的小规模环境(如最多4.8K状态、1.7K观测)的ATPO,RecBayes可在任意大的状态空间中运行,且不依赖状态或队友动作信息。在多智能体系统文献中的基准领域,经部分可观测化并扩展至100万状态和2^125种观测后,RecBayes仍能仅凭部分观测成功识别团队与任务,进而有效辅助团队完成任务。
原文摘要 · Abstract (English)
This paper proposes RecBayes, a novel approach for ad hoc teamwork under partial observability, a setting where agents are deployed on-the-fly to environments where pre-existing teams operate, that never requires, at any stage, access to the states of the environment or the actions of its teammates. We show that by relying on a recurrent Bayesian classifier trained using past experiences, an ad hoc agent is effectively able to identify known teams and tasks being performed from observations alone. Unlike recent approaches such as PO-GPL (Gu et al., 2021) and FEAT (Rahman et al., 2023), that require at some stage fully observable states of the environment, actions of teammates, or both, or approaches such as ATPO (Ribeiro et al., 2023) that require the environments to be small enough to be tabularly modelled (Ribeiro et al., 2023), in their work up to 4.8K states and 1.7K observations, we show RecBayes is both able to handle arbitrarily large spaces while never relying on either states and teammates' actions. Our results in benchmark domains from the multi-agent systems literature, adapted for partial observability and scaled up to 1M states and 2^125 observations, show that RecBayes is effective at identifying known teams and tasks being performed from partial observations alone, and as a result, is able to assist the teams in solving the tasks effectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。