arXiv:2409.11741cs.LGcs.AI2024-09ICRA被引 2

让非专家用最少干预指导多智能体动态分组协作。

HARP: Human-Assisted Regrouping with Permutation Invariant Critic for Multi-Agent Reinforcement Learning

  • 训练时智能体自动重组,部署时主动求助人类
  • 利用排列不变评论器评估人类建议,提升协作效率
  • 仅需少量非专家指导即可显著改善多智能体表现

人机协同强化学习通过引入人类专家经验加速智能体学习并在复杂任务中提供关键指导与反馈。然而,现有方法多聚焦于单智能体任务,且要求训练过程中持续的人类参与,大幅增加人力负担并限制可扩展性。本文提出 HARP(Human-Assisted Regrouping with Permutation Invariant Critic),一种面向群体任务的多智能体强化学习框架。该框架在训练阶段使智能体动态调整分组以优化协作,在部署阶段主动寻求人类协助,并利用排列不变的组别评论器评估与优化人类提出的分组方案,使非专家用户能以最小干预提供有效建议。在多个协作场景中,该方法仅需有限的非专家指导即可显著提升性能。项目代码见 https://github.com/huawen-hu/HARP。

原文摘要 · Abstract (English)

Human-in-the-loop reinforcement learning integrates human expertise to accelerate agent learning and provide critical guidance and feedback in complex fields. However, many existing approaches focus on single-agent tasks and require continuous human involvement during the training process, significantly increasing the human workload and limiting scalability. In this paper, we propose HARP (Human-Assisted Regrouping with Permutation Invariant Critic), a multi-agent reinforcement learning framework designed for group-oriented tasks. HARP integrates automatic agent regrouping with strategic human assistance during deployment, enabling and allowing non-experts to offer effective guidance with minimal intervention. During training, agents dynamically adjust their groupings to optimize collaborative task completion. When deployed, they actively seek human assistance and utilize the Permutation Invariant Group Critic to evaluate and refine human-proposed groupings, allowing non-expert users to contribute valuable suggestions. In multiple collaboration scenarios, our approach is able to leverage limited guidance from non-experts and enhance performance. The project can be found at https://github.com/huawen-hu/HARP.

多智能体人机协同动态分组

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。