用多模态模型预测空管指令与飞机动作的时间差和持续时间。
Multimodal Deep Learning for ATCO Command Lifecycle Modeling and Workload Prediction
- 融合结构化数据、轨迹序列和图像特征建模指令生命周期。
- 准确预测指令与飞机响应的时间偏移及持续时长,提升可解释性。
- 适合空管系统优化、智能指令生成与人员调度研究者参考。
空中交通管制员(ATCO)在高密度空域中发出高强度语音指令,精准的负荷建模对安全与效率至关重要。本文提出一种多模态深度学习框架,整合结构化数据、飞行轨迹序列与图像特征,以估计指令生命周期中的两个关键参数:指令与相应飞机机动之间的时间偏移,以及指令持续时间。构建了高质量数据集,通过滑动窗口与直方图法检测机动点。设计了CNN-Transformer集成模型,实现高精度、强泛化性和可解释性的预测。通过将飞行轨迹与语音指令关联,本工作首次提供支持智能指令生成的模型,具有实际应用价值,可用于负荷评估、人员配置与排班优化。
原文摘要 · Abstract (English)
Air traffic controllers (ATCOs) issue high-intensity voice commands in dense airspace, where accurate workload modeling is critical for safety and efficiency. This paper proposes a multimodal deep learning framework that integrates structured data, trajectory sequences, and image features to estimate two key parameters in the ATCO command lifecycle: the time offset between a command and the resulting aircraft maneuver, and the command duration. A high-quality dataset was constructed, with maneuver points detected using sliding window and histogram-based methods. A CNN-Transformer ensemble model was developed for accurate, generalizable, and interpretable predictions. By linking trajectories to voice commands, this work offers the first model of its kind to support intelligent command generation and provides practical value for workload assessment, staffing, and scheduling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。