arXiv:2603.11802cs.AI2026-03被引 2

提出半去中心化框架,解决通信不确定下的多智能体协作控制问题。

A Semi-Decentralized Approach to Multiagent Control

  • 引入半去中心化机制,允许智能体动态存储动作与观测
  • 构建SDec-POMDP模型,统一多种多智能体通信方式
  • 提出精确算法RS-SDA*,在多个基准和海上救援场景验证

我们提出一种表达性强的框架与算法,用于在通信不确定性环境下对协作智能体进行半去中心化控制。半马尔可夫控制允许动作执行时间分布,而半马尔可夫通信(即半去中心化)则允许智能体将动作与观测存储于历史中的时间分布。我们将半去中心化扩展至部分可观测马尔可夫决策过程(POMDP),构建出SDec-POMDP模型,该模型统一了去中心化与多智能体POMDP,并涵盖多种现有显式通信机制。本文提出递归小步半去中心化A*(RS-SDA*),一种生成SDec-POMDP最优策略的精确算法。在多个标准基准及海上医疗撤离场景的半去中心化版本上进行了评估。本研究为通过半去中心化视角探索各类多智能体通信问题提供了清晰的理论基础。

原文摘要 · Abstract (English)

We introduce an expressive framework and algorithms for the semi-decentralized control of cooperative agents in environments with communication uncertainty. Whereas semi-Markov control admits a distribution over time for agent actions, semi-Markov communication, or what we refer to as semi-decentralization, gives a distribution over time for what actions and observations agents can store in their histories. We extend semi-decentralization to the partially observable Markov decision process (POMDP). The resulting SDec-POMDP unifies decentralized and multiagent POMDPs and several existing explicit communication mechanisms. We present recursive small-step semi-decentralized A* (RS-SDA*), an exact algorithm for generating optimal SDec-POMDP policies. RS-SDA* is evaluated on semi-decentralized versions of several standard benchmarks and a maritime medical evacuation scenario. This paper provides a well-defined theoretical foundation for exploring many classes of multiagent communication problems through the lens of semi-decentralization.

多智能体强化学习决策优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。