arXiv:2605.04345cs.LG2026-05

揭示延迟环境下多智能体学习的结构等价性,为高效协同决策提供新路径。

Structural Equivalence and Learning Dynamics in Delayed MARL

  • 通过观测-动作历史证明观测延迟与动作延迟在策略集和轨迹分布上完全等价
  • 在非独立转移场景下,最小局部增强状态失效,学习动态出现根本差异
  • 利用等价性实现从观测延迟到动作延迟的零样本策略迁移,适合复杂延迟系统研究者

我们形式化证明了在合作部分可观测多智能体系统中,观测延迟(OD)与动作延迟(AD)通过观测-动作历史具有等价性。两者生成相同的可接受联合策略集合,且诱导的状态-动作-观测轨迹在分布上一致,从而在去中心化部分可观测马尔可夫决策过程(Dec-POMDPs)中产生相同的最优解。这一结果将原有的无限时域单智能体结论推广至任意时域、去中心化执行的多智能体部分可观测问题,并允许任何混合延迟配置简化为纯观测延迟系统。进一步,在转移独立马尔可夫决策过程(TI-MDPs)中,观测-动作历史可约简为可处理的最小局部增强状态。然而,数值实验表明,尽管最优解空间结构同构,实际学习动态却截然不同:第一,在非独立转移情况下,最小局部增强状态不再成立;第二,时序差分(TD)算法中的操作约束与因果信用分配误差导致不同延迟范式下的学习行为差异显著。最后,我们借助该结构等价性成功实现了从观测延迟到动作延迟的多智能体零样本策略迁移,为复杂延迟系统提供了统一高效的求解方法。

原文摘要 · Abstract (English)

We formally establish the equivalence between Observation Delay (OD) and Action Delay (AD) in cooperative partially observable multi-agent systems using observation-action histories. We show that both systems generate identical admissible joint-policy sets, and their induced state-action-observation trajectories are identical in distribution, leading to identical optimal solutions in Decentralized Partially Observable Markov Decision Processes (Dec-POMDPs). This formally generalizes existing infinite-horizon single-agent results to any-horizon partially observable cooperative multi-agent problems with decentralized policy execution, and allows any mixed-delay configuration to be reduced to a pure OD system. We further prove that in Transition-Independent MDPs (TI-MDPs), the observation-action history reduces to a tractable minimal local augmented state. However, we show through numerical experiments that although the optimal solution spaces are structurally isomorphic, the practical learning dynamics are fundamentally different. First, using the minimal local augmented state, the equivalence no longer holds when transitions are not independent. Second, operational constraints and causal credit-assignment errors in Temporal Difference (TD) algorithms induce different learning behaviors across regimes. Finally, leveraging this structural equivalence to bypass these learning challenges, we demonstrate successful multi-agent zero-shot policy transfer from OD to AD, paving the way for unified, efficient solution methods in complex delayed systems.

多智能体延迟学习策略迁移Dec-POMDP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。