用大模型分析飞行数据,自动识别非管制机场的空中安全隐患。
Towards Automated Air Traffic Safety Assessment Around Non-Towered Airports Using Large Language Models

- 融合语音、天气、航迹等多模态数据,用视觉语言模型分析飞行安全
- 仅用语音和天气数据,开源模型在危险判断任务上F1超0.85
- 适合航空安全研究者和智能交通系统开发者参考
本文研究利用大语言模型(LLMs)对非管制机场的飞行后安全进行评估。这类机场依赖通用交通咨询频率(CTAF)进行空中交通协调,因飞行员自主通报机制常发生近距碰撞。我们提出一种通用视觉语言模型(VLM)方法,分析转录的CTAF无线电通信、气象报告(METAR)、ADS-B飞行轨迹及目视飞行规则航图。以半月湾机场为例,通过真实飞行数据定性验证,并构建了一个包含12类风险的合成数据集,用于量化评估。该数据集涵盖语音与天气模态,测试了三个开源模型(Qwen 2.5-7B、Mistral-7B、Gemma-2-9B)和三个闭源模型(GPT-4o、GPT-5.4、Claude Sonnet 4.6)。即使仅使用CTAF和METAR输入,开源模型在二分类危险判定任务中平均宏F1超过0.85。未来工作将扩展至全模态定量评估与更多真实案例。结果表明,多模态大模型可成为非管制机场安全评估的重要工具。
原文摘要 · Abstract (English)
We investigate frameworks for post-flight safety analysis at non-towered airports using large language models (LLMs). Non-towered airports rely on the Common Traffic Advisory Frequency (CTAF) for air traffic coordination and experience frequent near mid-air collisions due to the pilot self-announcement communication protocol. We propose a general vision-language model (VLM) approach to analyze the transcribed CTAF radio communications in natural language, METeorological Aerodrome Report (METAR) weather data, Automatic Dependent Surveillance-Broadcast (ADS-B) flight trajectories, and Visual Flight Rules sectional charts of the airfield. We provide a preliminary study at Half Moon Bay Airport, with a qualitative real world case study and a quantitative evaluation using a new synthetic dataset of communications and weather modalities. We qualitatively evaluate our framework on real flight data using Gemini 2.5 Pro, demonstrating accurate identification of a right-of-way violation. The synthetic dataset is derived from real examples and includes a 12-category hazard taxonomy, and is used to benchmark three open-source (Qwen 2.5-7B, Mistral-7B, Gemma-2-9B) and three closed-source (GPT-4o, GPT-5.4, Claude Sonnet 4.6) LLM models on the subset of inputs related to CTAF and METAR. Even limited to CTAF and METAR inputs and open source LLMs, instances of our framework typically achieve a macro F1 score above 0.85 on a binary nominal/danger classification task. Future work includes a quantitative evaluation across all modalities and a larger number of real world examples. Taken together, our results suggest that VLM analysis of safety at non-towered airports may be a valuable future capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。