用扩散模型提升空管冲突检测与避让的决策灵活性。
Diffusion-RL Based Air Traffic Conflict Detection and Resolution Method
- 将扩散模型引入空管决策,生成多模式动作分布。
- 高密度场景下成功率94.1%,近地相撞减少59%。
- 适合需要高安全性和灵活应变的空管自动化系统。
在全球航空流量持续增长的背景下,高效且安全的冲突检测与规避(CD&R)对空中交通管理至关重要。尽管深度强化学习(DRL)为CD&R自动化提供了前景,但现有方法常因策略存在“单模态偏差”,在复杂动态约束下缺乏决策灵活性,易导致“决策死锁”。为此,本文首次将扩散概率模型融入安全关键的CD&R任务,提出新型自主冲突规避框架Diffusion-AC。与传统方法收敛至单一最优解不同,该框架将策略建模为由价值函数引导的逆去噪过程,能够生成丰富、高质量、多模态的动作分布。核心架构辅以密度渐进式安全课程(DPSC)训练机制,确保代理在从稀疏到高密度交通环境中的稳定高效学习。大量仿真实验表明,所提方法显著优于一系列先进DRL基准。尤其在最具挑战性的高密度场景中,Diffusion-AC不仅保持94.1%的高成功率,且相比次优基线近地相撞(NMAC)事件减少约59%,显著提升系统安全裕度。这一性能跃升源于其独特的多模态决策能力,使代理可灵活切换至有效替代机动方案。
原文摘要 · Abstract (English)
In the context of continuously rising global air traffic, efficient and safe Conflict Detection and Resolution (CD&R) is paramount for air traffic management. Although Deep Reinforcement Learning (DRL) offers a promising pathway for CD&R automation, existing approaches commonly suffer from a "unimodal bias" in their policies. This leads to a critical lack of decision-making flexibility when confronted with complex and dynamic constraints, often resulting in "decision deadlocks." To overcome this limitation, this paper pioneers the integration of diffusion probabilistic models into the safety-critical task of CD&R, proposing a novel autonomous conflict resolution framework named Diffusion-AC. Diverging from conventional methods that converge to a single optimal solution, our framework models its policy as a reverse denoising process guided by a value function, enabling it to generate a rich, high-quality, and multimodal action distribution. This core architecture is complemented by a Density-Progressive Safety Curriculum (DPSC), a training mechanism that ensures stable and efficient learning as the agent progresses from sparse to high-density traffic environments. Extensive simulation experiments demonstrate that the proposed method significantly outperforms a suite of state-of-the-art DRL benchmarks. Most critically, in the most challenging high-density scenarios, Diffusion-AC not only maintains a high success rate of 94.1% but also reduces the incidence of Near Mid-Air Collisions (NMACs) by approximately 59% compared to the next-best-performing baseline, significantly enhancing the system's safety margin. This performance leap stems from its unique multimodal decision-making capability, which allows the agent to flexibly switch to effective alternative maneuvers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。