arXiv:2410.13953cs.LG2024-10被引 6

用扩散模型解决多智能体部分可观测问题,实现状态重建与误差控制。

On Diffusion Models for Multi-Agent Partial Observability: Shared Attractors, Error Bounds, and Composite Flow

  • 基于局部历史构建扩散模型,将全局状态视为稳定不动点。
  • 发现近似误差与雅可比秩负相关,提出线性代理模型量化偏差。
  • 设计复合扩散过程,理论保证收敛到真实状态,适用于非共同可观测场景。

多智能体系统面临部分可观测性(PO)挑战,去中心化部分可观测马尔可夫决策过程(Dec-POMDP)揭示了这一问题的本质。尽管近期方法依赖深度学习模型应对PO,但对其如何影响智能体处理PO及交互行为的严格理解仍不足。本文研究在Dec-POMDP中,利用扩散模型从局部动作-观测历史重构全局状态。我们发现,以局部历史为条件的扩散模型将可能状态表示为稳定不动点:在可共同观测(CO)的Dec-POMDP中,各智能体的扩散模型共享唯一不动点对应全局状态;在非CO场景中,共享不动点则生成基于联合历史的状态分布。进一步发现,深度学习近似误差会导致不动点偏离真实状态,且该偏差与雅可比矩阵秩呈负相关。基于此低秩特性,我们构造一个代理线性回归模型来逼近扩散模型的局部行为,并由此建立偏差上界。基于该上界,提出一种 extit{复合扩散过程},通过智能体间迭代更新,在理论上收敛至真实状态。

原文摘要 · Abstract (English)

Multiagent systems grapple with partial observability (PO), and the decentralized POMDP (Dec-POMDP) model highlights the fundamental nature of this challenge. Whereas recent approaches to addressing PO have appealed to deep learning models, providing a rigorous understanding of how these models and their approximation errors affect agents' handling of PO and their interactions remain a challenge. In addressing this challenge, we investigate reconstructing global states from local action-observation histories in Dec-POMDPs using diffusion models. We first find that diffusion models conditioned on local history represent possible states as stable fixed points. In collectively observable (CO) Dec-POMDPs, individual diffusion models conditioned on agents' local histories share a unique fixed point corresponding to the global state, while in non-CO settings, shared fixed points yield a distribution of possible states given joint history. We further find that, with deep learning approximation errors, fixed points can deviate from true states and the deviation is negatively correlated to the Jacobian rank. Inspired by this low-rank property, we bound a deviation by constructing a surrogate linear regression model that approximates the local behavior of a diffusion model. With this bound, we propose a \emph{composite diffusion process} iterating over agents with theoretical convergence guarantees to the true state.

多智能体扩散模型状态估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。