测试发现多数协作强化学习基准无需真正解码部分可观测问题。
Probing Dec-POMDP Reasoning in Cooperative MARL
- 用统计与信息论方法诊断37个场景下的智能体决策复杂度。
- 超半数任务中反应式策略表现不输带记忆的智能体。
- 现有基准可能低估了协作难题,适合改进环境设计者看。
协作多智能体强化学习通常被建模为去中心化部分可观测马尔可夫决策过程(Dec-POMDP),其难点源于部分可观测性与去中心化协调。真正解决此类任务需具备Dec-POMDP推理能力,即通过历史推断隐藏状态并基于局部信息协调。然而,当前主流基准是否真正要求这种推理仍不明确。本文构建了一套诊断工具,结合统计性能对比与信息论探针,评估了IPPO和MAPPO在MPE、SMAX、Overcooked、Hanabi和MaBrax等37个场景中的行为复杂度。结果显示,在超过一半的场景中,反应式策略的表现与带记忆的智能体相当;涌现的协调往往依赖脆弱且同步的动作耦合,而非稳定的时序影响。这表明当前训练范式下,部分广泛使用的基准可能并未充分检验核心的Dec-POMDP假设,可能导致对进展的过度乐观评估。研究团队已公开诊断工具,以支持更严谨的环境设计与评估。
原文摘要 · Abstract (English)
Cooperative multi-agent reinforcement learning (MARL) is typically framed as a decentralised partially observable Markov decision process (Dec-POMDP), a setting whose hardness stems from two key challenges: partial observability and decentralised coordination. Genuinely solving such tasks requires Dec-POMDP reasoning, where agents use history to infer hidden states and coordinate based on local information. Yet it remains unclear whether popular benchmarks actually demand this reasoning or permit success via simpler strategies. We introduce a diagnostic suite combining statistically grounded performance comparisons and information-theoretic probes to audit the behavioural complexity of baseline policies (IPPO and MAPPO) across 37 scenarios spanning MPE, SMAX, Overcooked, Hanabi, and MaBrax. Our diagnostics reveal that success on these benchmarks rarely requires genuine Dec-POMDP reasoning. Reactive policies match the performance of memory-based agents in over half the scenarios, and emergent coordination frequently relies on brittle, synchronous action coupling rather than robust temporal influence. These findings suggest that some widely used benchmarks may not adequately test core Dec-POMDP assumptions under current training paradigms, potentially leading to over-optimistic assessments of progress. We release our diagnostic tooling to support more rigorous environment design and evaluation in cooperative MARL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。