arXiv:2511.02304cs.MAcs.AI2025-11被引 2

用有限状态机指导多智能体协作,实现高效多任务自主分工。

Automata-Conditioned Cooperative Multi-Agent Reinforcement Learning

  • 以自动机条件化策略,让智能体按任务状态动态调整行为。
  • 实验中实现按需分工的多步协作,如开门-守门-短路等复杂动作。
  • 支持新任务无需重训,适合需要灵活调度的多智能体系统。

我们研究在集中训练、分散执行框架下,学习多任务、多智能体协同策略。通过使用自动机表示分配给智能体的任务,可将团队目标分解为更简单的子任务。然而,现有方法样本效率低,且仅限于单任务场景,每新增任务均需重新训练策略。本文提出自动机条件化协同多智能体强化学习(ACC-MARL),一种学习任务条件化、去中心化的团队策略框架。我们识别了实现该框架的挑战,提出相应解决方案,并证明了所提方法的最优性。进一步表明,学习到的价值函数可用于测试时最优任务分配。实验显示,智能体间涌现出任务感知的多步协作行为,例如按按钮解锁门、持门不闭、短接任务等。

原文摘要 · Abstract (English)

We study learning multi-task, multi-agent policies for cooperative, temporal objectives, under centralized training, decentralized execution. In this setting, using automata to represent tasks assigned to agents enables breaking down a team-level objective into simpler, smaller sub-tasks. However, existing approaches remain sample-inefficient and are limited to the single-task case, requiring retraining policies for each new task. In this work, we present Automata-Conditioned Cooperative Multi-Agent Reinforcement Learning (ACC-MARL), a framework for learning task-conditioned, decentralized team policies. We identify challenges to the feasibility of ACC-MARL, propose solutions, and prove that our approach is optimal. We further show that learned value functions can be used to assign tasks optimally at test time. Experiments demonstrate emergent task-aware, multi-step coordination among agents, such as pressing a button to unlock a door, holding the door, and short-circuiting tasks.

多智能体自动机协作强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。