arXiv:2507.10142cs.AIcs.LG2025-07综述被引 1

提出可适应性框架,系统梳理MARL在真实场景中的假设失效问题。

Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review

  • 以假设感知视角划分学习、策略、场景三类适应性维度。
  • 指出现有评估缺乏对分布或结构变化的明确定义与诊断能力。
  • 适合关注MARL落地挑战与评估方法的研究者阅读。

多智能体强化学习(MARL)在模拟基准中表现强劲,但实际部署常违背算法设计与评估所依赖的假设。智能体数量可能变动,目标可能改变,中心化信息可能不可用,执行可能异步,合作者策略可能未知。现有综述讨论了可扩展性、鲁棒性、泛化性与可迁移性等理想属性,但这些术语常指代不同分析对象与不同类型的分布或结构变化。本文提出‘可适应性’作为假设感知的分类框架,而非要求所有MARL算法在所有场景下都成功。区分三个维度:学习适应性(训练或系统假设改变时学习范式的适用性)、策略适应性(部署时对已学策略的复用或调整)、场景驱动适应性(基准与评估协议是否暴露可控且具诊断意义的变化)。通过明确变化内容、发生时机、允许的适应方式及成功标准,该框架厘清了既有概念的关系,并指出现有MARL评估仍存在定义模糊之处。

原文摘要 · Abstract (English)

Multi-Agent Reinforcement Learning (MARL) has achieved strong performance in simulated benchmarks, yet real deployments often violate the assumptions under which algorithms are designed and evaluated. Agent populations may change, objectives may shift, centralized information may be unavailable, execution may become asynchronous, and partner policies may be unfamiliar. Existing surveys discuss related desiderata such as scalability, robustness, generalization, and transferability, but these terms often refer to different objects of analysis and different kinds of distributional or structural shift. This survey proposes \textit{adaptability} as an assumption-aware taxonomy for organizing these shifts, rather than as a universal requirement that every MARL algorithm should succeed in every setting. We distinguish three dimensions: \textit{learning adaptability}, which concerns the applicability of learning paradigms under changed training or system assumptions; \textit{policy adaptability}, which concerns the reuse or adaptation of learned policies under deployment-time changes; and \textit{scenario-driven adaptability}, which concerns whether benchmarks and evaluation protocols expose controlled, diagnostically useful shifts. By separating what changes, when the change occurs, what adaptation is allowed, and what success means, the framework clarifies how established concepts fit together and identifies where current MARL evaluation remains underspecified.

多智能体强化学习评估框架适应性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。