针对复杂三维防御场景,提出可扩展的智能体协同强化学习框架。
Embedded Mean Field Reinforcement Learning for Perimeter-defense Game
- 用均值场方法实现高阶动作聚合,支持大规模防御者协同
- 在多尺度仿真中显著提升收敛速度与整体表现,优于现有基线
- 结合轻量注意力机制,适合真实复杂环境中的实时决策
随着无人机和导弹技术的快速发展,关键区域的攻防围城博弈日益复杂且战略意义重大。然而,现有研究多局限于小规模、简化的二维场景,常忽略真实环境扰动、运动动力学及固有异质性等关键因素,限制了实际应用。为此,本文研究三维空间下具有异质性的大规模围城防御博弈,引入运动动力学与风场等真实要素,推导攻防双方纳什均衡策略,刻画胜利区域,并通过大量仿真实验验证理论成果。为应对大规模异质控制挑战,提出嵌入式均值场演员-评论家(EMFAC)框架:利用表示学习实现均值场层面的高层动作聚合,支持防御者可扩展协同;同时引入基于奖励表示的轻量级代理注意力机制,选择性过滤观测与均值场信息,提升决策效率并加速收敛。多尺度仿真实验表明,EMFAC在收敛速度与整体性能上均优于现有基线。进一步在小规模真实世界实验中测试并分析,验证其在复杂场景下的实用性。
原文摘要 · Abstract (English)
With the rapid advancement of unmanned aerial vehicles (UAVs) and missile technologies, perimeter-defense game between attackers and defenders for the protection of critical regions have become increasingly complex and strategically significant across a wide range of domains. However, existing studies predominantly focus on small-scale, simplified two-dimensional scenarios, often overlooking realistic environmental perturbations, motion dynamics, and inherent heterogeneity--factors that pose substantial challenges to real-world applicability. To bridge this gap, we investigate large-scale heterogeneous perimeter-defense game in a three-dimensional setting, incorporating realistic elements such as motion dynamics and wind fields. We derive the Nash equilibrium strategies for both attackers and defenders, characterize the victory regions, and validate our theoretical findings through extensive simulations. To tackle large-scale heterogeneous control challenges in defense strategies, we propose an Embedded Mean-Field Actor-Critic (EMFAC) framework. EMFAC leverages representation learning to enable high-level action aggregation in a mean-field manner, supporting scalable coordination among defenders. Furthermore, we introduce a lightweight agent-level attention mechanism based on reward representation, which selectively filters observations and mean-field information to enhance decision-making efficiency and accelerate convergence in large-scale tasks. Extensive simulations across varying scales demonstrate the effectiveness and adaptability of EMFAC, which outperforms established baselines in both convergence speed and overall performance. To further validate practicality, we test EMFAC in small-scale real-world experiments and conduct detailed analyses, offering deeper insights into the framework's effectiveness in complex scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。