arXiv:2608.15372cs.AIcs.MA2026-08

让无人机群在通信中断下自适应对抗,通过三重机制提升实战鲁棒性。

UC-PSRO: Utility-Conditioned Policy-Space Response Oracles with a Communication-Dropout Curriculum for Game-Theoretic Course-of-Action Generation in Adversarial Swarms

  • 用自对弈强化学习让蓝红双方互为最佳应对策略,避免固定对手陷阱。
  • 通信丢包率升至75%时任务成功率反升至62%,证明去中心化设计有效。
  • 单次推理可调参适配不同作战意图,适合快速部署的军事智能系统。

我们研究在通信受限环境下,针对自适应红方对手,生成蓝方无人机集群的游戏理论最优行动方案(COA),灵感来自美国空军公开的SBIR招标。提出UC-PSRO方法,融合三项机制:(i) PSRO自对弈训练,使蓝红策略互为近似最优响应,而非一方对抗固定脚本;(ii) 使用FiLM对指挥官意图向量进行条件控制,从狄利克雷分布采样,在不重新训练的情况下实现执行时策略重定向;(iii) 训练中渐进式丢弃通信图边,促使集群学习去中心化、同侪互助的容错机制。在模拟的未分类海事场景中评估,设定N=25个蓝方智能体,共5组随机种子,并扩展至N=200。结果发现真实权衡:仅通信丢包课程即带来最强最稳健的任务完成率,随丢包率从0升至0.75,成功率反而从35%升至62%;加入效用条件与自对弈训练虽能减缓收敛速度,但未显著优于固定对手策略,两者在统计上无差异且差距极小。我们如实报告该收敛代价尚未被鲁棒性收益抵消,拒绝夸大单一方法优势,并提供可在单块消费级显卡上以毫秒级每步完成200智能体训练的完全向量化开源环境。

原文摘要 · Abstract (English)

We study generating game-theoretically optimized Courses of Action (COAs) for a Blue UAS swarm against an adaptive Red adversary in a communication-degraded environment, motivated by (but not derived from) a public U.S. Air Force SBIR solicitation. We propose UC-PSRO (Utility-Conditioned Policy-Space Response Oracles with a Communication-Dropout Curriculum), combining three mechanisms: (i) PSRO self-play, so Blue and Red policies train as approximate best responses to each other rather than one side against a fixed scripted opponent; (ii) FiLM conditioning of the Blue policy on a Commander's-Intent weight vector, sampled from a Dirichlet distribution during training, so one trained policy is re-steerable at execution time without retraining; and (iii) a curriculum annealing communication-graph edge dropout during training, so the swarm learns decentralized, peer-to-peer fallback instead of depending on full connectivity. We evaluate on a synthetic, unclassified stand-in for the solicitation's maritime scenario, with 5 seeds at N=25 Blue agents and a scalability sweep to N=200. We find a genuine trade-off, not a uniform win: the communication-dropout curriculum alone gives the strongest, most robust mission-completion rates of any learned method, improving counter-intuitively as denial increases (35% to 62% success as dropout rises from 0 to 0.75); adding utility-conditioning and PSRO self-play substantially slows convergence within a fixed budget, and we find no reliable exploitability advantage for self-play over a fixed-opponent policy, both statistically indistinguishable from a small, near-zero gap. We report this honestly as a convergence cost not yet offset by a demonstrated robustness benefit, rather than overstating one method as dominant, and provide a fully vectorized, open environment training at N=200 agents in single-digit milliseconds per step on a single consumer GPU.

博弈论无人机群通信容错自对弈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。