arXiv:2506.21127cs.LGcs.AI2025-06

用自适应策略切换提升无人机在对抗空域的导航安全

Meta Policy Switching for Secure UAV Deconfliction in Adversarial Airspace

  • 构建多抗扰策略库,通过动态选择应对未知攻击
  • 在3D复杂环境中实现90%以上无冲突轨迹率
  • 适合高安全要求的无人机自主系统研发者

基于强化学习的自主无人机导航易受对抗攻击影响,导致行为异常与任务失败。现有鲁棒强化学习方法因依赖固定扰动设定,难以泛化到未见或分布外(OOD)攻击。为此,本文提出元策略切换框架,通过元策略动态选择多个鲁棒策略以应对未知对抗扰动。核心为折扣版汤普森采样(DTS)机制,将策略选择建模为多臂赌博机问题,利用自生成对抗观测最小化价值分布偏移。首先在不同扰动强度下训练多样化的动作鲁棒策略集合;DTS元策略在线自适应选择,优化对自诱导、分段平稳攻击的鲁棒性。理论分析表明,该机制可最小化期望遗憾,实现对OOD攻击的自适应鲁棒性,并表现出涌现的抗脆弱特性。在包含复杂3D障碍物的场景中,针对白盒(投影梯度下降)和黑盒(GPS欺骗)攻击的大量仿真验证显示,本方法显著提升导航效率与无冲突轨迹率,优于标准鲁棒与基础强化学习基线,凸显其实际安全与可靠性优势。

原文摘要 · Abstract (English)

Autonomous UAV navigation using reinforcement learning (RL) is vulnerable to adversarial attacks that manipulate sensor inputs, potentially leading to unsafe behavior and mission failure. Although robust RL methods provide partial protection, they often struggle to generalize to unseen or out-of-distribution (OOD) attacks due to their reliance on fixed perturbation settings. To address this limitation, we propose a meta-policy switching framework in which a meta-level polic dynamically selects among multiple robust policies to counter unknown adversarial shifts. At the core of this framework lies a discounted Thompson sampling (DTS) mechanism that formulates policy selection as a multi-armed bandit problem, thereby minimizing value distribution shifts via self-induced adversarial observations. We first construct a diverse ensemble of action-robust policies trained under varying perturbation intensities. The DTS-based meta-policy then adaptively selects among these policies online, optimizing resilience against self-induced, piecewise-stationary attacks. Theoretical analysis shows that the DTS mechanism minimizes expected regret, ensuring adaptive robustness to OOD attacks and exhibiting emergent antifragile behavior under uncertainty. Extensive simulations in complex 3D obstacle environments under both white-box (Projected Gradient Descent) and black-box (GPS spoofing) attacks demonstrate significantly improved navigation efficiency and higher conflict free trajectory rates compared to standard robust and vanilla RL baselines, highlighting the practical security and dependability benefits of the proposed approach.

无人机对抗攻击强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。