arXiv:2602.16196cs.LGcs.AI2026-02被引 2

通过采样强关联智能体,实现异构多智能体强化学习的高效协作。

Graphon Mean-Field Subsampling for Cooperative Heterogeneous Multi-Agent Reinforcement Learning

  • 基于图论权重采样关键智能体,近似异构交互的平均场
  • 样本复杂度为κ的多项式,最优差距达O(1/√κ)
  • 适用于大规模机器人协同场景,理论与仿真验证有效

在多智能体强化学习中,协调大量交互智能体是核心挑战,联合状态-动作空间随智能体数量呈指数增长。平均场方法通过聚合交互缓解此问题,但假设所有交互同质。近期图论框架能捕捉异质性,但随着智能体增多计算开销巨大。为此,我们提出GMFS框架——一种图论平均场子采样方法,用于可扩展的异构多智能体协作强化学习。通过按交互强度采样κ个智能体,近似图论加权平均场,学习策略的样本复杂度为poly(κ),最优差距为O(1/√κ)。我们在机器人协同任务中进行数值模拟验证理论,结果表明GMFS可达到近似最优性能。

原文摘要 · Abstract (English)

Coordinating large populations of interacting agents is a central challenge in multi-agent reinforcement learning (MARL), where the size of the joint state-action space scales exponentially with the number of agents. Mean-field methods alleviate this burden by aggregating agent interactions, but these approaches assume homogeneous interactions. Recent graphon-based frameworks capture heterogeneity, but are computationally expensive as the number of agents grows. Therefore, we introduce $\texttt{GMFS}$, a $\textbf{G}$raphon $\textbf{M}$ean-$\textbf{F}$ield $\textbf{S}$ubsampling framework for scalable cooperative MARL with heterogeneous agent interactions. By subsampling $κ$ agents according to interaction strength, we approximate the graphon-weighted mean-field and learn a policy with sample complexity $\mathrm{poly}(κ)$ and optimality gap $O(1/\sqrtκ)$. We verify our theory with numerical simulations in robotic coordination, showing that $\texttt{GMFS}$ achieves near-optimal performance.

多智能体平均场图论强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。