arXiv:2411.04976cs.LG2024-11

让智能体在信息不一致时仍能协作,打破传统假设的局限。

Noisy Zero-Shot Coordination: Breaking The Common Knowledge Assumption In Zero-Shot Coordination Games

  • 构建带噪声观测的元决策模型,模拟真实世界信息差异。
  • 在多任务训练中学习适应不同噪声版本,实现强泛化协作能力。
  • 适合复杂现实场景下需要动态协作的强化学习应用。

零样本协作(ZSC)是研究强化学习智能体与新伙伴协作能力的热门设定。以往的ZSC假设问题设置为共同知识:每个智能体都知道底层的Dec-POMDP,并且知道对方也知晓这一点,以此类推。然而,在复杂现实环境中,这种假设往往不成立,因为环境难以被完整且正确地定义。因此,当共同知识假设失效时,传统的ZSC训练方法可能无法有效协作。为解决此问题,我们提出了噪声零样本协作(NZSC)问题:智能体观测到的是地面真实Dec-POMDP的不同噪声版本,这些噪声版本服从固定的噪声模型。仅地面真实Dec-POMDP的分布和噪声模型为共同知识。我们证明,通过构建包含所有真实Dec-POMDP的扩展状态空间,可将一个NZSC问题转化为标准的ZSC问题。针对求解NZSC问题,我们提出一种简单灵活的元学习方法——NZSC训练,即智能体在一系列协调问题上进行训练,但只能观察到其噪声版本。实验表明,经过NZSC训练的强化学习智能体即使在精确问题设置并非共同知识的情况下,也能与新伙伴良好协作。

原文摘要 · Abstract (English)

Zero-shot coordination (ZSC) is a popular setting for studying the ability of reinforcement learning (RL) agents to coordinate with novel partners. Prior ZSC formulations assume the $\textit{problem setting}$ is common knowledge: each agent knows the underlying Dec-POMDP, knows others have this knowledge, and so on ad infinitum. However, this assumption rarely holds in complex real-world settings, which are often difficult to fully and correctly specify. Hence, in settings where this common knowledge assumption is invalid, agents trained using ZSC methods may not be able to coordinate well. To address this limitation, we formulate the $\textit{noisy zero-shot coordination}$ (NZSC) problem. In NZSC, agents observe different noisy versions of the ground truth Dec-POMDP, which are assumed to be distributed according to a fixed noise model. Only the distribution of ground truth Dec-POMDPs and the noise model are common knowledge. We show that a NZSC problem can be reduced to a ZSC problem by designing a meta-Dec-POMDP with an augmented state space consisting of all the ground-truth Dec-POMDPs. For solving NZSC problems, we propose a simple and flexible meta-learning method called NZSC training, in which the agents are trained across a distribution of coordination problems - which they only get to observe noisy versions of. We show that with NZSC training, RL agents can be trained to coordinate well with novel partners even when the (exact) problem setting of the coordination is not common knowledge.

零样本协作强化学习元学习分布式决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。