用超性质逻辑指导部分可观测多智能体强化学习,提升策略表达力。
HyPOLE: Hyperproperty-Guided Multi-Agent Reinforcement Learning under Partial Observation

- 基于超时逻辑HyperLTL设计可表达复杂约束的学习目标
- 在SMAC、MessySMAC和WildFire上显著优于基线方法
- 适合需要形式化约束的多智能体协同任务研究者
形式化规范是引导学习过程的强大工具,相较于奖励塑形具有三大优势:数学严谨性、对目标与约束的表达能力,以及定义达成目标策略的能力。然而这些优势在多智能体强化学习(MARL)中尚未充分探索。本文提出HyPOLE框架,用于部分可观测环境下的MARL,通过超性质(hyperproperties)特别是超时逻辑HyperLTL来指导学习。结合集中训练分散执行(CTDE)技术,合成去中心化策略。在SMAC、MessySMAC和WildFire基准上的评估表明,该方法显著优于现有基线。
原文摘要 · Abstract (English)
Formal specification is a powerful tool to guide the learning process and provides significant advantages over reward shaping: (1) mathematical rigor; (2) expressiveness to specify objectives and constraints, and (3) the ability to define tactics to achieve objectives. However, these benefits remain largely unexplored in the context of Multi-Agent Reinforcement Learning (MARL). This paper introduces HyPOLE, a novel framework for MARL under partial observability, where learning is guided by the expressive power of the so-called hyperproperties and, in particular, the temporal logic HyperLTL. We integrate Centralized Training for Decentralized Execution (CTDE) techniques with HyPOLE to synthesize decentralized policies, and our evaluation on SMAC, MessySMAC, and WildFire benchmark demonstrates clear advantages over baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。