arXiv:2510.20218cs.MAcs.AI2025-10NeurIPS被引 1

提出线性复杂度的交互建模框架,让多智能体协作更高效且可解释。

High-order Interactions Modeling for Interpretable Multi-Agent Q-Learning

  • 用连分数结构建模任意阶交互,计算复杂度仅为O(n)
  • 在MPE和StarCraft场景中性能优于基线,且可解释性更强
  • 适合需要理解协作机制的多智能体系统研究者

在多智能体强化学习中,建模智能体间交互对有效协作与理解合作机制至关重要。然而,以往高阶交互建模方法常受组合爆炸或黑箱网络结构的限制。本文提出一种新型价值分解框架——连分数Q学习(QCoFr),能以仅O(n)的线性复杂度灵活捕捉任意阶智能体交互,避免建模丰富协作时的组合爆炸问题。此外,引入变分信息瓶颈提取潜在信息用于信用分配估计,帮助智能体过滤噪声交互,显著提升协作能力与可解释性。大量实验表明,QCoFr不仅持续取得更优性能,其可解释性也与理论分析一致。

原文摘要 · Abstract (English)

The ability to model interactions among agents is crucial for effective coordination and understanding their cooperation mechanisms in multi-agent reinforcement learning (MARL). However, previous efforts to model high-order interactions have been primarily hindered by the combinatorial explosion or the opaque nature of their black-box network structures. In this paper, we propose a novel value decomposition framework, called Continued Fraction Q-Learning (QCoFr), which can flexibly capture arbitrary-order agent interactions with only linear complexity $\mathcal{O}\left({n}\right)$ in the number of agents, thus avoiding the combinatorial explosion when modeling rich cooperation. Furthermore, we introduce the variational information bottleneck to extract latent information for estimating credits. This latent information helps agents filter out noisy interactions, thereby significantly enhancing both cooperation and interpretability. Extensive experiments demonstrate that QCoFr not only consistently achieves better performance but also provides interpretability that aligns with our theoretical analysis.

多智能体强化学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。