arXiv:2607.20547cs.LG2026-07

用强化学习让无人机在信号差时自动避让,保障空中交通安全。

Conflict Resolution under Degraded Surveillance in Air Corridors Using Multi-Agent Reinforcement Learning

论文配图:Conflict Resolution under Degraded Surveillance in Air Corridors Using Multi-Agent Reinforcement Learning
图 1 · 摘自论文原文
  • 分角色训练无人机和电动垂直起降飞行器的独立避障策略。
  • 90组测试中冲突解决率超99%,多数在1秒内完成。
  • 适合研究空中交通管理与低可靠通信环境下的飞行控制。

安全的先进空中交通管理(AAM)要求飞行器在监视信息噪声大、延迟、不完整或临时不可用时仍保持间距。本研究提出基于深度Q网络的多智能体强化学习框架,用于在三维结构化空走廊中,对异构的小型无人航空器与电动垂直起降飞机进行去中心化冲突解决。为两类飞行器分别训练独立策略,基于局部观测和包含14种动作(保持航向、转向、垂直机动、着陆、速度控制)的行动空间。仿真融合了飞行器动力学特性、能耗、走廊约束、观测噪声、通信延迟、信息丢失、风扰动、执行器不确定性及模型不确定性。在90种交通密度与最小间距阈值组合下评估训练策略。分离失效频率与持续时间随交通密度和间距要求上升,但大多数冲突可在1秒内解决。安全状态下,智能体保持运动约79%时间;冲突期间,转向占33%,保持航向29%,速度控制25%,垂直机动13%。六组帕累托最优配置揭示了安全与走廊容量间的权衡。该框架支持在监视条件退化场景下对更安全的AAM冲突解决策略进行仿真评估。

原文摘要 · Abstract (English)

Safe Advanced Air Mobility operations require aircraft to maintain separation when surveillance information is noisy, delayed, incomplete, or temporarily unavailable. This study develops a Deep Q-Network-based Multi-Agent Reinforcement Learning framework for decentralized conflict resolution among heterogeneous small unmanned aerial vehicles and electric vertical takeoff and landing aircraft operating within a structured three-dimensional corridor. Separate policies are trained for the two aircraft categories using local observations and a 14-action space that includes maintaining course, turning, vertical maneuvering, landing, and speed control. The simulation incorporates aircraft-specific dynamics, energy use, corridor constraints, observation noise, communication delay, information dropout, wind disturbance, actuator uncertainty, and model uncertainty. The trained policies are evaluated across 90 combinations of traffic density and minimum separation thresholds. Loss-of-separation frequency and duration generally increase with traffic density and separation requirements, although most events are resolved within 1s. Under safe conditions, agents maintain their motion approximately 79% of the time. During conflicts, turning accounts for 33% of actions, followed by maintaining motion at 29%, speed control at 25%, and vertical maneuvers at 13%. Six Pareto-optimal configurations reveal trade-offs between safety and corridor capacity. The framework supports the simulation-based evaluation of safer AAM conflict-resolution strategies under degraded surveillance conditions.

空中交通强化学习无人机避障

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。