arXiv:2502.17618cs.LGcs.AI2025-02中稿 · AAMAS 2025被引 6

从多样演示中学习团队协作的多模式行为,提升复杂任务下的协同效率。

Hierarchical Imitation Learning of Team Behavior from Heterogeneous Demonstrations

  • 构建分层策略框架,分因子学习异构演示中的团队行为
  • 在多种协作场景中超越基线方法,实现高精度行为建模
  • 适合需要多模式协作的多智能体系统与人机协同研究

成功协作需团队成员保持同步,尤其在复杂的序列任务中。成员需动态协调执行子任务的顺序与方式。然而,现实约束如部分可观测性和通信带宽有限常导致协作不佳。即使专家团队,同一任务也可能有多种执行方式。为开发此类任务的多智能体系统与人机协作,我们关注从数据中学习多模式团队行为。多智能体模仿学习(MAIL)提供了一种从演示中数据驱动学习团队行为的可行框架,但现有方法难以处理异构演示,因它们假设所有演示来自单一团队策略。本文提出DTIL:一种用于复杂序列任务中学习多模式团队行为的分层MAIL算法。DTIL通过分层策略表示每个团队成员,并以因子化方式从异构团队演示中学习这些策略。采用分布匹配方法,有效缓解误差累积,可扩展至长时程和连续状态表示。实验表明,DTIL优于现有MAIL基线,在多种协作场景中准确建模团队行为。

原文摘要 · Abstract (English)

Successful collaboration requires team members to stay aligned, especially in complex sequential tasks. Team members must dynamically coordinate which subtasks to perform and in what order. However, real-world constraints like partial observability and limited communication bandwidth often lead to suboptimal collaboration. Even among expert teams, the same task can be executed in multiple ways. To develop multi-agent systems and human-AI teams for such tasks, we are interested in data-driven learning of multimodal team behaviors. Multi-Agent Imitation Learning (MAIL) provides a promising framework for data-driven learning of team behavior from demonstrations, but existing methods struggle with heterogeneous demonstrations, as they assume that all demonstrations originate from a single team policy. Hence, in this work, we introduce DTIL: a hierarchical MAIL algorithm designed to learn multimodal team behaviors in complex sequential tasks. DTIL represents each team member with a hierarchical policy and learns these policies from heterogeneous team demonstrations in a factored manner. By employing a distribution-matching approach, DTIL mitigates compounding errors and scales effectively to long horizons and continuous state representations. Experimental results show that DTIL outperforms MAIL baselines and accurately models team behavior across a variety of collaborative scenarios.

多智能体模仿学习团队协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。