arXiv:2506.04215cs.MAcs.AI2025-06被引 2

提出新型策略框架,让多智能体在视野有限时仍能接近最优决策。

Thinking Beyond Visibility: A Near-Optimal Policy Framework for Locally Interdependent Multi-Agent MDPs

  • 构建可记忆超视野信息的扩展截断策略,突破传统方法局限。
  • 在小固定视野下显著减少停滞现象,性能提升明显。
  • 适用于协作导航等场景,尤其适合对稳定性要求高的应用。

去中心化部分可观测马尔可夫决策过程(Dec-POMDP)被证明是NEXP-完全问题,难以求解。然而,在协同导航、避障和编队控制等问题中,可假设局部可见性与局部依赖性。DeWeese和Qu(2024)提出了局部相互依赖多智能体MDP模型,并给出三种闭式可计算的策略,其性能在各种情况下均指数级接近最优,且与可见性相关。但研究发现,当可见性较小且固定时,这些策略表现不佳,常因“惩罚抖动”现象陷入模拟停滞。本文提出扩展截断策略类(Extended Cutoff Policy Class),据我们所知,这是首个在任意局部相互依赖多智能体MDP中,与可见性呈指数级接近最优的非平凡闭式部分可观测策略类。该策略能记忆超出自身视野的信息,显著提升小固定视野下的性能,解决惩罚抖动问题,并在特定条件下保证联合最优行为。此外,还推广了原模型以支持转移依赖和扩展奖励依赖,理论结果在新设定中依然成立。

原文摘要 · Abstract (English)

Decentralized Partially Observable Markov Decision Processes (Dec-POMDPs) are known to be NEXP-Complete and intractable to solve. However, for problems such as cooperative navigation, obstacle avoidance, and formation control, basic assumptions can be made about local visibility and local dependencies. The work DeWeese and Qu 2024 formalized these assumptions in the construction of the Locally Interdependent Multi-Agent MDP. In this setting, it establishes three closed-form policies that are tractable to compute in various situations and are exponentially close to optimal with respect to visibility. However, it is also shown that these solutions can have poor performance when the visibility is small and fixed, often getting stuck during simulations due to the so called "Penalty Jittering" phenomenon. In this work, we establish the Extended Cutoff Policy Class which is, to the best of our knowledge, the first non-trivial class of near optimal closed-form partially observable policies that are exponentially close to optimal with respect to the visibility for any Locally Interdependent Multi-Agent MDP. These policies are able to remember agents beyond their visibilities which allows them to perform significantly better in many small and fixed visibility settings, resolve Penalty Jittering occurrences, and under certain circumstances guarantee fully observable joint optimal behavior despite the partial observability. We also propose a generalized form of the Locally Interdependent Multi-Agent MDP that allows for transition dependence and extended reward dependence, then replicate our theoretical results in this setting.

多智能体决策优化策略设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。