用强化学习优化无人机群防御中的拦截优先级,提升保护关键区域的效率。
Reinforcement Learning for Decision-Level Interception Prioritization in Drone Swarm Defense
- 基于强化学习的决策代理在离散动作空间中选择最优拦截目标。
- 相比人工规则基线,平均损伤降低,防御效率显著提升。
- 适用于需快速战略决策的复杂防御系统,可无缝集成现有体系。
低成本自杀式无人机群威胁日益严峻,对现代防御系统提出快速且战略性决策要求,需在多个效应器和高价值目标区之间合理分配拦截任务。本文通过案例研究展示强化学习在此挑战中的实际优势。我们构建了一个高保真仿真环境,模拟真实作战约束条件,训练一个决策层强化学习代理,协调多个效应器实现最优拦截优先级分配。该代理在离散动作空间中,根据位置、类别及效应器状态等观测特征,决定每个效应器应拦截的目标。在数百个模拟攻击场景中,与手工设计的规则基线对比,强化学习策略始终实现更低的平均损伤和更高的防御效率。本案例表明,强化学习可作为防御架构中的战略层,增强系统韧性而不替换现有控制系统。所有代码与仿真资源均已公开,视频演示展示了策略的定性行为。
原文摘要 · Abstract (English)
The growing threat of low-cost kamikaze drone swarms poses a critical challenge to modern defense systems demanding rapid and strategic decision-making to prioritize interceptions across multiple effectors and high-value target zones. In this work, we present a case study demonstrating the practical advantages of reinforcement learning in addressing this challenge. We introduce a high-fidelity simulation environment that captures realistic operational constraints, within which a decision-level reinforcement learning agent learns to coordinate multiple effectors for optimal interception prioritization. Operating in a discrete action space, the agent selects which drone to engage per effector based on observed state features such as positions, classes, and effector status. We evaluate the learned policy against a handcrafted rule-based baseline across hundreds of simulated attack scenarios. The reinforcement learning based policy consistently achieves lower average damage and higher defensive efficiency in protecting critical zones. This case study highlights the potential of reinforcement learning as a strategic layer within defense architectures, enhancing resilience without displacing existing control systems. All code and simulation assets are publicly released for full reproducibility, and a video demonstration illustrates the policy's qualitative behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。