arXiv:2412.15639cs.MAcs.AI2024-12中稿 · AAMAS 2025被引 2

让智能体不靠通信也能默契配合,自适应筛选关键信息。

Tacit Learning with Adaptive Information Selection for Cooperative Multi-Agent Reinforcement Learning

  • 通过自适应信息选择机制,让智能体自主判断哪些信息重要。
  • 在无通信环境下,仍能实现接近有通信的协作效果。
  • 适合通信受限或需低延迟决策的多智能体系统。

在多智能体强化学习(MARL)中,集中训练、分散执行(CTDE)框架因表现优异而被广泛采用。然而,其进一步发展面临两大挑战:一是智能体难以自主评估输入信息对协作任务的相关性,影响决策能力;二是通信受限且环境部分可观测时,智能体无法获取全局信息,限制了从全局视角协同的能力。为此,本文提出一种基于信息选择与隐式学习的新型协作MARL框架。该框架使智能体在训练过程中逐步形成隐性协调机制,无需通信即可在离散空间中推断其他智能体的协作行为,仅依赖局部信息完成高效协作。同时,引入门控与选择机制,使智能体能根据环境变化动态过滤信息,提升决策能力。在主流MARL基准上的实验表明,该框架可无缝集成现有先进算法,并带来显著性能提升。

原文摘要 · Abstract (English)

In multi-agent reinforcement learning (MARL), the centralized training with decentralized execution (CTDE) framework has gained widespread adoption due to its strong performance. However, the further development of CTDE faces two key challenges. First, agents struggle to autonomously assess the relevance of input information for cooperative tasks, impairing their decision-making abilities. Second, in communication-limited scenarios with partial observability, agents are unable to access global information, restricting their ability to collaborate effectively from a global perspective. To address these challenges, we introduce a novel cooperative MARL framework based on information selection and tacit learning. In this framework, agents gradually develop implicit coordination during training, enabling them to infer the cooperative behavior of others in a discrete space without communication, relying solely on local information. Moreover, we integrate gating and selection mechanisms, allowing agents to adaptively filter information based on environmental changes, thereby enhancing their decision-making capabilities. Experiments on popular MARL benchmarks show that our framework can be seamlessly integrated with state-of-the-art algorithms, leading to significant performance improvements.

多智能体强化学习信息选择隐式协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。