现有协作MARL方法依赖脆弱约定而非真实推理,新环境需强化观测与记忆机制。
Remembering the Markov Property in Cooperative MARL
- 用简单循环网络学习合作惯例,绕过真实环境信号
- 同模型在需真实推理时可学出稳健策略,证明能力非局限
- 适合关注协作智能、环境设计的科研人员
协作多智能体强化学习(MARL)通常形式化为去中心化部分可观测马尔可夫决策过程(Dec-POMDP),要求智能体基于部分信息推断环境与其他智能体的行为。当前无模型MARL算法普遍采用简单的循环函数逼近器来应对这一挑战。本文认为,这些方法的实证成功并非源于有效恢复马尔可夫信号,而是通过学习绕过环境观测与记忆的简单协作惯例。通过针对性案例研究,我们发现共适应智能体可习得脆弱的约定,当与非自适应智能体配对时即失效。关键的是,相同模型在任务设计强制要求时能学习到基于实际观测的稳健策略,表明问题不在于学习模型本身,而在于基准环境设计不足。分析还暗示现代MARL环境可能未能充分检验Dec-POMDP的核心假设。因此,我们倡导构建以两个核心原则为基础的新协作环境:(1) 行为基于观测;(2) 基于记忆推理其他智能体,确保成功依赖真实技能而非脆弱的共适应协议。
原文摘要 · Abstract (English)
Cooperative multi-agent reinforcement learning (MARL) is typically formalised as a Decentralised Partially Observable Markov Decision Process (Dec-POMDP), where agents must reason about the environment and other agents' behaviour. In practice, current model-free MARL algorithms use simple recurrent function approximators to address the challenge of reasoning about others using partial information. In this position paper, we argue that the empirical success of these methods is not due to effective Markov signal recovery, but rather to learning simple conventions that bypass environment observations and memory. Through a targeted case study, we show that co-adapting agents can learn brittle conventions, which then fail when partnered with non-adaptive agents. Crucially, the same models can learn grounded policies when the task design necessitates it, revealing that the issue is not a fundamental limitation of the learning models but a failure of the benchmark design. Our analysis also suggests that modern MARL environments may not adequately test the core assumptions of Dec-POMDPs. We therefore advocate for new cooperative environments built upon two core principles: (1) behaviours grounded in observations and (2) memory-based reasoning about other agents, ensuring success requires genuine skill rather than fragile, co-adapted agreements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。