arXiv:2510.07813cs.AI2025-10

智能体在追逃中权衡通信暴露风险与信息获取,提升追捕成功率。

Strategic Communication under Threat: Learning Information Trade-offs in Pursuit-Evasion Games

  • 设计多头强化学习框架,融合导航、通信与对手行为预测
  • 实验显示其追捕成功率超越6个基线方法
  • 对通信风险和能力差异具有强泛化能力,适合军事仿真应用

对抗环境中,智能体需在获取信息与暴露自身之间权衡。为此,我们构建了追逃-暴露-隐藏博弈(PEEC),其中追捕者需决定何时通信以获取逃逸者位置,但每次通信会暴露自身位置,增加被攻击风险。双方通过强化学习训练运动策略,追捕者还学习通信策略以平衡可见性与风险。本文提出SHADOW框架,集成连续导航控制、离散通信动作及对手建模,实现多头序列决策。实验证明,采用SHADOW的追捕者成功率显著高于六个竞争基线。消融实验表明时序建模与对手建模对决策至关重要。敏感性分析显示,所学策略在不同通信风险及双方物理能力差异下均具有良好泛化性能。

原文摘要 · Abstract (English)

Adversarial environments require agents to navigate a key strategic trade-off: acquiring information enhances situational awareness, but may simultaneously expose them to threats. To investigate this tension, we formulate a PursuitEvasion-Exposure-Concealment Game (PEEC) in which a pursuer agent must decide when to communicate in order to obtain the evader's position. Each communication reveals the pursuer's location, increasing the risk of being targeted. Both agents learn their movement policies via reinforcement learning, while the pursuer additionally learns a communication policy that balances observability and risk. We propose SHADOW (Strategic-communication Hybrid Action Decision-making under partial Observation for Warfare), a multi-headed sequential reinforcement learning framework that integrates continuous navigation control, discrete communication actions, and opponent modeling for behavior prediction. Empirical evaluations show that SHADOW pursuers achieve higher success rates than six competitive baselines. Our ablation study confirms that temporal sequence modeling and opponent modeling are critical for effective decision-making. Finally, our sensitivity analysis reveals that the learned policies generalize well across varying communication risks and physical asymmetries between agents.

强化学习博弈论追逃游戏通信策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。