arXiv:2605.17393cs.AIcs.LG2026-05

提出新型协作图结构,让多智能体通信更高效且有理论保障。

Heterogeneous Information-Bottleneck Coordination Graphs for Multi-Agent Reinforcement Learning

论文配图:Heterogeneous Information-Bottleneck Coordination Graphs for Multi-Agent Reinforcement Learning
图 1 · 摘自论文原文
  • 基于信息瓶颈构建分组对齐的稀疏图,决定连边存在与密度。
  • 每组内可差异化控制通信容量,压缩消息保留关键内容。
  • 理论证明可优化拓扑学习,适合复杂协作任务中的智能体设计。

协作图是合作式多智能体强化学习的核心抽象,但现有稀疏图学习方法缺乏理论依据来决定哪些边应存在以及每条边应承载多少信息。当前方法依赖启发式准则,无法保证学习到的拓扑结构合理性,也无原则性方法为不同结构关系分配差异化的通信容量。为此,我们提出异质信息瓶颈协作图(HIBCG),在群组感知的稀疏图中同时确定边的存在性与消息容量,具有理论依据。通过图信息瓶颈(GIB)作为基础工具,HIBCG首先构建分组对齐的块对角先验,提供闭式标准以决定哪些边应保留及每组块的边密度;随后在所得拓扑上控制各智能体特征带宽,压缩消息仅保留任务相关部分。我们证明:分组对齐先验严格收紧了拓扑学习的变分界,目标可按组块分解,实现差异化边控制,且容量分配遵循水灌原则。

原文摘要 · Abstract (English)

Coordination graphs are a central abstraction in cooperative multi-agent reinforcement learning (MARL), yet existing sparse-graph learners lack a theoretically grounded mechanism to decide which edges should exist and how much information each edge should carry. Current methods rely on heuristic criteria that offer no formal guarantee on the learned topology, and no principled way to allocate different communication capacities to structurally different agent relationships. To address this, we propose Heterogeneous Information-Bottleneck Coordination Graphs (HIBCG), which learns a group-aware sparse graph in which both edge existence and message capacity are theoretically justified. With the graph information bottleneck (GIB) serving as the underlying tool, HIBCG first constructs a group-aligned block-diagonal prior that provides a closed-form criterion for edge retention -- determining which edges should exist and at what density per group block -- and then controls per-agent feature bandwidth on the resulting topology, compressing messages to retain only task-relevant content. We prove that the group-aligned prior strictly tightens the variational bound on topology learning, that the objective decomposes per group block, enabling differential edge control, and that capacity allocation follows a water-filling principle.

多智能体信息瓶颈协作图强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。