arXiv:2503.15703cs.MAcs.AI2025-03被引 3

用任务并行性判断多智能体是否该分工,理论可预测性能提升。

Predicting Multi-Agent Specialization via Task Parallelizability

  • 基于任务并行度和团队规模,给出性能提升的闭式预测公式。
  • 在SMAC和MPE环境中,理论预测与实际分工效果高度一致。
  • 可诊断强化学习算法缺陷,指导智能体策略设计。

在多智能体系统中,何时应鼓励分工,何时应训练通用型智能体独立完成任务?我们提出,分工效果主要取决于任务的并行性——即多个智能体同时执行任务组件的潜力。受分布式系统中Amdahl定律启发,我们提出了一个仅依赖任务并发性和团队规模的闭式边界,用于预测分工能否提升性能。我们在两个标准多智能体强化学习(MARL)基准上验证该模型:星战多智能体挑战(SMAC,无限并发)和多粒子环境(MPE,单位容量瓶颈),发现理论边界在两种极端情况下均与实证的分工程度高度吻合。三项后续实验在过厨房-AI环境中进一步验证,该模型适用于存在复杂空间与资源瓶颈、策略多样化的场景。除预测外,该边界还可作为诊断工具,揭示当前MARL训练算法在更大状态空间下收敛至专用策略时的偏差。

原文摘要 · Abstract (English)

When should we encourage specialization in multi-agent systems versus train generalists that perform the entire task independently? We propose that specialization largely depends on task parallelizability: the potential for multiple agents to execute task components concurrently. Drawing inspiration from Amdahl's Law in distributed systems, we present a closed-form bound that predicts when specialization improves performance, depending only on task concurrency and team size. We validate our model on two standard MARL benchmarks that represent opposite regimes -- StarCraft Multi-Agent Challenge (SMAC, unlimited concurrency) and Multi-Particle Environment (MPE, unit-capacity bottlenecks) -- and observe close alignment between the bound at each extreme and an empirical measure of specialization. Three follow-up experiments in Overcooked-AI demonstrate that the model works in environments with more complex spatial and resource bottlenecks that allow for a range of strategies. Beyond prediction, the bound also serves as a diagnostic tool, highlighting biases in MARL training algorithms that cause sub-optimal convergence to specialist strategies with larger state spaces.

多智能体分工预测强化学习并行性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。