arXiv:2602.14559cs.LGcs.AI2026-02

让智能体可动态创建同伴,实现自适应团队规模。

Fluid-Agent Reinforcement Learning

  • 设计可自动生成新智能体的流式环境框架
  • 团队规模随环境需求动态调整,提升适应性
  • 适用于需要灵活协作的现实场景如生物分裂、企业分拆

多智能体强化学习(MARL)通常研究固定数量智能体间的互动。然而现实中,智能体数量往往不固定且未知,且智能体可自主创建新个体(如细胞分裂或公司分立)。本文提出允许智能体生成其他智能体的框架,称为流式智能体环境。我们引入博弈论解概念以应对此类动态结构,并在多种基准任务中评估了多个MARL算法的表现,包括可动态产生智能体的猎物-捕食者和基于层级觅食任务。此外,我们引入了一个新环境,展示流式机制如何催生传统固定种群设置下无法出现的新策略。实验表明,该框架能生成根据环境需求自动调节规模的智能体团队。

原文摘要 · Abstract (English)

The primary focus of multi-agent reinforcement learning (MARL) has been to study interactions among a fixed number of agents embedded in an environment. However, in the real world, the number of agents is neither fixed nor known a priori. Moreover, an agent can decide to create other agents (for example, a cell may divide, or a company may spin off a division). In this paper, we propose a framework that allows agents to create other agents; we call this a fluid-agent environment. We present game-theoretic solution concepts for fluid-agent games and empirically evaluate the performance of several MARL algorithms within this framework. Our experiments include fluid variants of established benchmarks such as Predator-Prey and Level-Based Foraging, where agents can dynamically spawn, as well as a new environment we introduce that highlights how fluidity can unlock novel solution strategies beyond those observed in fixed-population settings. We demonstrate that this framework yields agent teams that adjust their size dynamically to match environmental demands.

多智能体强化学习动态协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。