arXiv:2511.17367cs.LG2025-11

提出首个部分可观测下鲁棒实时追捕策略,解决真实场景追捕难题。

R2PS: Worst-Case Robust Real-Time Pursuit Strategies under Partial Observability

  • 基于信念保持机制扩展动态规划,适应追捕者信息不全场景。
  • 在未见图结构上实现零样本泛化,性能优于现有方法。
  • 适合安全监控、无人机追捕等需实时决策的场景。

在部分可观测条件下计算追捕博弈(PEGs)的最坏情况鲁棒策略耗时严重。尽管强化学习(RL)方法如等价策略泛化(EPG)和Grasper能学习对不同游戏动态鲁棒的图神经网络(GNN)策略,但它们仅适用于完全信息场景,且未考虑逃逸者预测追捕者动作的可能性。本文首次提出部分可观测下的最坏情况鲁棒实时追捕策略(R2PS)。我们证明了传统动态规划(DP)算法在逃逸者异步行动下仍保持最优性;随后设计了关于逃逸者可能位置的信念保持机制,将DP策略拓展至部分可观测环境;最后将该机制嵌入EPG框架,通过跨图强化学习对抗异步移动的DP逃逸策略,构建出可实时执行的追捕策略。训练后,该策略在未见过的真实图结构上表现出鲁棒的零样本泛化能力,且持续优于现有游戏强化学习方法直接在测试图上训练的策略。

原文摘要 · Abstract (English)

Computing worst-case robust strategies in pursuit-evasion games (PEGs) is time-consuming, especially when real-world factors like partial observability are considered. While important for general security purposes, real-time applicable pursuit strategies for graph-based PEGs are currently missing when the pursuers only have imperfect information about the evader's position. Although state-of-the-art reinforcement learning (RL) methods like Equilibrium Policy Generalization (EPG) and Grasper provide guidelines for learning graph neural network (GNN) policies robust to different game dynamics, they are restricted to the scenario of perfect information and do not take into account the possible case where the evader can predict the pursuers' actions. This paper introduces the first approach to worst-case robust real-time pursuit strategies (R2PS) under partial observability. We first prove that a traditional dynamic programming (DP) algorithm for solving Markov PEGs maintains optimality under the asynchronous moves by the evader. Then, we propose a belief preservation mechanism about the evader's possible positions, extending the DP pursuit strategies to a partially observable setting. Finally, we embed the belief preservation into the state-of-the-art EPG framework to finish our R2PS learning scheme, which leads to a real-time pursuer policy through cross-graph reinforcement learning against the asynchronous-move DP evasion strategies. After reinforcement learning, our policy achieves robust zero-shot generalization to unseen real-world graph structures and consistently outperforms the policy directly trained on the test graphs by the existing game RL approach.

追捕博弈强化学习部分可观测实时策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。