arXiv:2603.28900cs.ROcs.AI2026-03被引 1

用强化学习让小型无人机在GPS受干扰时仍能安全避撞。

Robust Multi-Agent Reinforcement Learning for Small UAS Separation Assurance under GPS Degradation and Spoofing

  • 设计对抗性训练机制,实时生成最危险的信号干扰。
  • 在35%干扰率下碰撞率接近零,性能远超普通方法。
  • 适合高密度无人机编队、抗欺骗场景下的安全控制。

针对小型无人机(sUAS)在GPS信号退化或被欺骗情况下的分离保障问题,本文采用多智能体强化学习(MARL)方法。在协同监视中,各飞行器广播基于GPS的位置信息;当这些位置信息被篡改时,整体空域状态变得不可靠。我们将此状态观测污染建模为智能体与对手之间的零和博弈:以概率R,对手对观测状态施加最大危害性扰动以降低各智能体的安全性能。本文推导出该对抗扰动的闭式表达,避免了传统对抗训练中的迭代优化,实现状态维度线性时间评估。该表达近似于在建模不确定性集上价值函数的精确最小值,精度达二阶。进一步证明,在Kullback-Leibler正则化下,干净与受污染观测间的安全性能差距最多随干扰概率线性恶化。最后,将闭式对抗策略集成至MARL策略梯度算法中,获得鲁棒应对策略。在高密度sUAS仿真中,当干扰水平高达35%时,碰撞率几乎为零,显著优于未考虑对抗扰动的基线策略。

原文摘要 · Abstract (English)

We address robust separation assurance for small Unmanned Aircraft Systems (sUAS) under GPS degradation and spoofing via Multi-Agent Reinforcement Learning (MARL). In cooperative surveillance, each aircraft (or agent) broadcasts its GPS-derived position; when such position broadcasts are corrupted, the entire observed air traffic state becomes unreliable. We cast this state observation corruption as a zero-sum game between the agents and an adversary: with probability R, the adversary perturbs the observed state to maximally degrade each agent's safety performance. We derive a closed-form expression for this adversarial perturbation, bypassing the iterative inner optimization of adversarial training entirely and enabling linear-time evaluation in the state dimension. We show that this expression approximates the exact minimizer of the value function over the modeled uncertainty set with second-order accuracy. We further bound the safety performance gap between clean and corrupted observations, showing that it degrades at most linearly with the corruption probability under Kullback-Leibler regularization. Finally, we integrate the closed-form adversarial policy into a MARL policy gradient algorithm to obtain a robust counter-policy for the agents. In a high-density sUAS simulation, we observe near-zero collision rates under corruption levels up to 35%, outperforming a baseline policy trained without adversarial perturbations.

无人机强化学习安全控制抗欺骗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。