arXiv:2504.04675cs.AIcs.LO2025-04NeurIPS被引 5

用超时序逻辑指导多智能体强化学习,提升复杂任务的控制策略性能。

HypRL: Reinforcement Learning of Control Policies for Hyperproperties

  • 通过斯科伦化处理超时序逻辑中的量词交替,构建可优化的奖励函数。
  • 在安全规划、深海宝藏等任务中,显著提高目标公式的满足概率。
  • 适合需要形式化约束的多智能体系统设计,如自动驾驶与分布式控制。

多智能体强化学习在复杂任务中的奖励设计仍面临挑战,现有方法常无法找到最优解或效率低下。本文提出HYPRL,一种基于超时序逻辑(HyperLTL)规范的强化学习框架,用于学习满足超性质的控制策略。通过斯科伦化处理量化符交替,并定义定量鲁棒性函数,在马尔可夫决策过程(转移未知)的执行轨迹上塑造奖励。选用合适的强化学习算法,使策略联合最大化期望奖励,从而提升超性质公式ϕ的满足概率。我们在包括安全感知规划、Deep Sea Treasure及贴现对应问题在内的多个基准任务上评估HYPRL,与基于规范的基线方法对比,验证了其有效性和高效性。

原文摘要 · Abstract (English)

Reward shaping in multi-agent reinforcement learning (MARL) for complex tasks remains a significant challenge. Existing approaches often fail to find optimal solutions or cannot efficiently handle such tasks. We propose HYPRL, a specification-guided reinforcement learning framework that learns control policies w.r.t. hyperproperties expressed in HyperLTL. Hyperproperties constitute a powerful formalism for specifying objectives and constraints over sets of execution traces across agents. To learn policies that maximize the satisfaction of a HyperLTL formula $ϕ$, we apply Skolemization to manage quantifier alternations and define quantitative robustness functions to shape rewards over execution traces of a Markov decision process with unknown transitions. A suitable RL algorithm is then used to learn policies that collectively maximize the expected reward and, consequently, increase the probability of satisfying $ϕ$. We evaluate HYPRL on a diverse set of benchmarks, including safety-aware planning, Deep Sea Treasure, and the Post Correspondence Problem. We also compare with specification-driven baselines to demonstrate the effectiveness and efficiency of HYPRL.

强化学习超时序逻辑多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。