arXiv:2603.15907cs.LGcs.SY2026-03

用博弈论加速强化学习,提前终止训练提升防御效率

Game-Theory-Assisted Reinforcement Learning for Border Defense: Early Termination based on Analytical Solutions

  • 用阿波罗尼斯圆计算追击阶段均衡,避免学习追击策略
  • 在单/多防御场景下奖励提升10%-20%,收敛更快
  • 适合复杂对抗环境下的强化学习训练优化

博弈论为对抗行为提供最优性保障,但当假设如完全信息不成立时,其性能会退化。强化学习虽具适应性,但在大规模复杂环境中样本效率低。本文提出一种混合方法,利用博弈论洞察提升强化学习训练效率。研究有限感知范围下的边境防御游戏,防御者表现依赖搜索与追击策略,经典微分博弈解法失效。方法采用阿波罗尼斯圆(Apollonius Circle, AC)计算检测后的均衡状态,实现强化学习回合的早期终止,无需学习追击动态。使强化学习聚焦于学习搜索策略,同时保证检测后行为最优。在单/多防御者设置中,该方法使奖励提升10%-20%,收敛速度加快,搜索轨迹更高效。大量实验验证了该方法的有效性。

原文摘要 · Abstract (English)

Game theory provides the gold standard for analyzing adversarial engagements, offering strong optimality guarantees. However, these guarantees often become brittle when assumptions such as perfect information are violated. Reinforcement learning (RL), by contrast, is adaptive but can be sample-inefficient in large, complex domains. This paper introduces a hybrid approach that leverages game-theoretic insights to improve RL training efficiency. We study a border defense game with limited perceptual range, where defender performance depends on both search and pursuit strategies, making classical differential game solutions inapplicable. Our method employs the Apollonius Circle (AC) to compute equilibrium in the post-detection phase, enabling early termination of RL episodes without learning pursuit dynamics. This allows RL to concentrate on learning search strategies while guaranteeing optimal continuation after detection. Across single- and multi-defender settings, this early termination method yields 10-20% higher rewards, faster convergence, and more efficient search trajectories. Extensive experiments validate these findings and demonstrate the overall effectiveness of our approach.

强化学习博弈论边境防御早期终止

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。