arXiv:2603.03664eess.SYcs.LG2026-03被引 1

提出可计算的通信学习框架,让多智能体在部分可观测环境下高效协作。

Principled Learning-to-Communicate with Quasi-Classical Information Structures

  • 基于信息结构理论,将通信学习问题分类为可计算与不可计算两类。
  • 设计新条件确保通信后仍保持可计算性,避免复杂度爆炸。
  • 给出可证明的算法,适用于实际场景中的多智能体协同决策。

在部分可观测环境中,深度多智能体强化学习中的学习通信(LTC)受到越来越多关注,其中控制与通信策略联合学习。同时,通信对决策的影响在控制理论中已有广泛研究。本文通过信息结构(IS)的视角,将这两条研究线统一起来,形式化了去中心化部分可观测马尔可夫决策过程(Dec-POMDPs)中的LTC问题,并基于信息共享前的IS对LTC问题进行分类。我们首先证明,一般情况下非经典LTC是计算上不可行的,因此聚焦于准经典(QC)LTC。接着提出一系列条件,满足这些条件时,通信后仍保持准经典信息结构;违反则可能导致普遍的计算困难。进一步,我们开发出可证明的规划与学习算法,并为满足上述条件的若干典型例子建立了准多项式时间与样本复杂度。此外,还揭示了严格准经典信息结构与策略无关共同信念(SI-CIBs)之间的关系,并解决了无需计算不可行预言机但超出具有SI-CIBs的去中心化马尔可夫决策过程的新问题,其结果可能具有独立意义。

原文摘要 · Abstract (English)

Learning-to-communicate (LTC) in partially observable environments has received increasing attention in deep multi-agent reinforcement learning, where the control and communication strategies are jointly learned. Meanwhile, the impact of communication on decision-making has been extensively studied in control theory. In this paper, we seek to formalize and better understand LTC by bridging these two lines of work, through the lens of information structures (ISs). To this end, we formalize LTC in decentralized partially observable Markov decision processes (Dec-POMDPs) under the common-information-based framework from decentralized stochastic control, and classify LTC problems based on the ISs before (additional) information sharing. We first show that non-classical LTCs are computationally intractable in general, and thus focus on quasi-classical (QC) LTCs. We then propose a series of conditions for QC LTCs, under which LTC preserves the QC IS after information sharing, whereas violating them can cause computational hardness in general. Further, we develop provable planning and learning algorithms for QC LTCs, and establish quasi-polynomial time and sample complexities for several QC LTC examples that satisfy the above conditions. Along the way, we also establish new results on a relationship between (strictly) QC IS and the condition of having strategy-independent common-information-based beliefs (SI-CIBs), as well as on solving Dec-POMDPs without computationally intractable oracles but beyond those with SI-CIBs, which may be of independent interest.

多智能体通信学习信息结构决策优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。