arXiv:2412.15700cs.AIcs.LG2024-12AAAI

提出一种统一个体与集体探索的自适应强化学习方法

AIR: Unifying Individual and Collective Exploration in Cooperative Multi-Agent Reinforcement Learning

  • 通过轨迹身份识别与动作选择对抗机制实现自适应探索
  • 理论证明可同时促进个体与集体探索,提升训练效率
  • 适用于多智能体协作场景,尤其适合复杂任务建模

合作式多智能体强化学习中的探索问题因缺乏显式策略而持续面临挑战。现有方法分别依赖于基于系统不确定性的个体探索和通过智能体间行为多样性实现的集体探索,但引入额外结构常导致训练效率下降且难以融合。本文提出自适应身份识别探索(AIR),包含两个对抗组件:从轨迹中识别智能体身份的分类器,以及根据状态自适应调整探索模式与程度的动作选择器。理论上证明AIR可在训练中同时促进个体与集体探索;实验表明其在多种任务上兼具高效性与有效性。

原文摘要 · Abstract (English)

Exploration in cooperative multi-agent reinforcement learning (MARL) remains challenging for value-based agents due to the absence of an explicit policy. Existing approaches include individual exploration based on uncertainty towards the system and collective exploration through behavioral diversity among agents. However, the introduction of additional structures often leads to reduced training efficiency and infeasible integration of these methods. In this paper, we propose Adaptive exploration via Identity Recognition~(AIR), which consists of two adversarial components: a classifier that recognizes agent identities from their trajectories, and an action selector that adaptively adjusts the mode and degree of exploration. We theoretically prove that AIR can facilitate both individual and collective exploration during training, and experiments also demonstrate the efficiency and effectiveness of AIR across various tasks.

多智能体强化学习探索策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。