arXiv:2509.09752cs.SDcs.CY2025-09

用语音和文字特征识别飞行员指令,提升无塔台机场管理效率

Combining Textual and Spectral Features for Robust Classification of Pilot Communications

  • 双通道模型:同时分析语音频谱和文字内容
  • F1分数超91%,深度模型结合频谱特征效果最佳
  • 无需新增设备,适合普通通用航空机场部署

准确估计飞机起飞与着陆等运行状态对机场管理至关重要,尤其在缺乏专用监视设施的非塔台机场仍具挑战。本文提出一种新型双通道机器学习框架,通过融合文本与频谱特征对飞行员无线电通信进行分类。基于美国某非塔台机场采集的真实音频数据,由持证飞行员标注操作意图,并经自动语音识别与梅尔频谱图提取预处理。我们评估了多种传统分类器及深度学习模型(包括集成方法、LSTM、CNN),在双通道管道中进行对比。据我们所知,这是首个在真实空管音频上使用双通道机器学习框架实现飞行操作意图分类的系统。结果表明,结合频谱特征与深度架构的模型表现最优,F1分数超过91%。数据增强进一步提升了对实际语音变化的鲁棒性。该方法具备可扩展性、低成本与无需额外基础设施的特点,为通用航空机场提供实用的空中交通监控解决方案。

原文摘要 · Abstract (English)

Accurate estimation of aircraft operations, such as takeoffs and landings, is critical for effective airport management, yet remains challenging, especially at non-towered facilities lacking dedicated surveillance infrastructure. This paper presents a novel dual pipeline machine learning framework that classifies pilot radio communications using both textual and spectral features. Audio data collected from a non-towered U.S. airport was annotated by certified pilots with operational intent labels and preprocessed through automatic speech recognition and Mel-spectrogram extraction. We evaluate a wide range of traditional classifiers and deep learning models, including ensemble methods, LSTM, and CNN across both pipelines. To our knowledge, this is the first system to classify operational aircraft intent using a dual-pipeline ML framework on real-world air traffic audio. Our results demonstrate that spectral features combined with deep architectures consistently yield superior classification performance, with F1-scores exceeding 91%. Data augmentation further improves robustness to real-world audio variability. The proposed approach is scalable, cost-effective, and deployable without additional infrastructure, offering a practical solution for air traffic monitoring at general aviation airports.

语音识别空管系统双通道模型分类任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。