arXiv:2605.08754cs.AI2026-05

提出新强化学习框架,实现实时多飞机滑行路径安全调度。

Value-Decomposed Reinforcement Learning Framework for Taxiway Routing with Hierarchical Conflict-Aware Observations

论文配图:Value-Decomposed Reinforcement Learning Framework for Taxiway Routing with Hierarchical Conflict-Aware Observations
图 1 · 摘自论文原文
  • 构建分层交通预判表征,捕捉当前及后续冲突信息。
  • 在长沙黄花机场测试中,安全与效率平衡优于基线方法。
  • 适合需要实时决策的空管自动化系统研究者参考。

滑行道路径规划与地面冲突避让是机场地面对运行中的安全关键决策问题。现有规划与优化方法常受限于在线计算成本,而强化学习方法可能难以有效表示下游交通冲突并平衡多重目标。本文提出冲突感知滑行路径(CaTR)框架,用于实时多飞机滑行路径规划。CaTR采用基于网格的机场表面环境与动作掩码机制,引入分层前瞻性交通表征以编码当前及后续冲突相关交通状态,并采用价值分解强化学习策略优先处理稀疏但关键的安全目标。在基于长沙黄花国际机场的真实环境中,于多种交通密度水平下进行实验。结果表明,CaTR在安全-效率权衡上优于代表性规划、优化及强化学习基线方法,同时保持实用的运行时间。

原文摘要 · Abstract (English)

Taxiway routing and on-surface conflict avoidance are coupled safety-critical decision problems in airport surface operations. Existing planning and optimization methods are often limited by online computational cost, while reinforcement learning methods may struggle to represent downstream traffic conflicts and balance multiple objectives. This paper presents Conflict-aware Taxiway Routing (CaTR), a reinforcement learning framework for real-time multi-aircraft taxiway routing. CaTR constructs a grid-based airport surface environment with action masking, introduces a hierarchical foresight traffic representation to encode current and downstream conflict-related traffic conditions, and adopts a value-decomposed reinforcement learning strategy to prioritize sparse but safety-critical objectives. Experiments are conducted on a realistic environment based on Changsha Huanghua International Airport under multiple traffic density levels. Results show that CaTR achieves better safety--efficiency trade-offs than representative planning, optimization, and reinforcement learning baselines while maintaining practical runtime.

强化学习滑行调度多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。