arXiv:2506.21129cs.LGcs.AI2025-06被引 1

让无人机在欺骗攻击下仍能稳定飞行,通过渐进式训练提升抗干扰能力。

Curriculum-Adapted Robust Reinforcement Learning for UAV Deconfliction in Adversarial Environments

  • 设计渐进式对抗训练流程,逐步增强扰动强度并保持误差分布一致。
  • 在固定和动态欺骗攻击下,任务成功率提升至近100%,基线仅20-56%。
  • 适合需高鲁棒性的自主飞行系统,尤其对抗未知攻击场景。

自主无人机日益依赖强化学习进行导航。然而,全球导航卫星系统(GNSS)欺骗攻击可引发分布外观测偏移,扭曲价值估计并降低任务性能。现有鲁棒强化学习方法通常仅对特定攻击模型有效,难以泛化到训练中未见的攻击。为此,本文提出一种课程引导的自适应框架,逐步将鲁棒策略暴露于强度递增的梯度攻击观测扰动中,同时保持各阶段时序差分(TD)误差分布一致。该方法不针对特定攻击模型,而是通过维持TD误差一致性来促进跨攻击条件的迁移能力。我们进一步推导出一个TD空间泛化保证:若测试时攻击引起的TD误差分布与最终课程阶段足够接近,则性能下降被可控。该框架在含动态3D障碍物的无人机避障环境中评估,面对先前未见过的固定与动态欺骗攻击。在固定欺骗条件下,课程自适应策略任务成功率接近100%,而标准与鲁棒强化学习基线仅为20-56%;在动态诱骗障碍物攻击下,获得最高每回合奖励,并在高空中交通密度下减少最多45%的任务完成步数。

原文摘要 · Abstract (English)

Autonomous unmanned aerial vehicles (UAVs) increasingly rely on reinforcement learning (RL) for navigation. However, global navigation satellite system (GNSS) spoofing attacks can induce out-of-distribution observation shifts that corrupt value estimation and degrade mission performance. Existing robust RL approaches typically improve resilience against specific attack models but often fail to generalize to attacks not encountered during training. To address this limitation, we propose a curriculum-guided adaptation framework that progressively exposes a robust policy to gradient-based adversarial observation perturbations of increasing intensity while aligning temporal-difference (TD) error distributions across curriculum stages. Rather than adapting to a particular attack model, the proposed approach preserves TD-error consistency to promote transferability across attack conditions. We further derive a TD-space generalization certificate showing that if the TD-error distribution induced by a test-time attack remains sufficiently close to that of the final curriculum stage, the resulting performance degradation is bounded. The framework is evaluated in a UAV deconfliction environment with dynamic 3D obstacles under previously unseen fixed and dynamic GNSS spoofing attacks. Under fixed spoofing conditions, the curriculum-adapted policy achieved near-perfect mission success rates, compared with 20-56% for standard and robust RL baselines. Under dynamic obstacle-luring spoofing attacks, it achieved the highest episodic rewards while reducing mission completion steps by up to 45% across increasing aerial traffic densities.

强化学习无人机对抗攻击鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。