为航空管制语言系统设计安全导向评估框架,揭示大模型在高危错误上的脆弱性。
Safety-Oriented Evaluation of Language Understanding Systems for Air Traffic Control

- 构建后果感知的评估框架,区分错误影响等级
- 峰值风险得分仅0.69,远低于安全部署要求
- 高危实体出错集中,暴露模型结构缺陷
航空管制(ATC)是典型的安全关键领域,指令误读可能导致严重操作后果。尽管大语言模型(LLMs)表现出较强的通用性能,其在实际ATC环境中的可靠性仍不明确。现有评估方法多依赖F1或宏准确率等整体指标,对所有错误一视同仁,未能考虑高风险语义错误(如跑道编号或移动约束错误)的不对称后果。为此,我们提出一种面向安全、后果感知的ATC评估框架。结果显示,尽管当前模型在整体准确率上表现尚可,但其实际运行可靠性严重受限:在干净转录文本上,最高风险得分为0.69,多数模型得分低于0.6,即使宏F1较高。进一步分析表明,错误集中在高影响实体,而动作类型分类相对稳定,反映出模型在结构性语义理解上的不足。这些发现凸显了在部署AI辅助航空管制系统前,必须采用后果感知的评估协议。
原文摘要 · Abstract (English)
Air Traffic Control (ATC) is a safety-critical domain in which incorrect interpretation of instructions may lead to severe operational consequences. While large language models (LLMs) demonstrate strong general performance, their reliability in operational ATC environments remains unclear. Existing evaluation approaches, largely based on aggregate metrics such as F1 or macro accuracy, treat all errors uniformly and fail to account for the asymmetric consequences of high-risk semantic mistakes (e.g., incorrect runway identifiers or movement constraints). To address this gap, we propose a safety-oriented, consequence-aware evaluation framework tailored to ATC operations. Our results reveal that while current LLMs achieve reasonable aggregate accuracy, their operational reliability is severely limited. Evaluated on clean transcripts, the peak Risk Score reaches only 0.69, with most models scoring below 0.6 despite high macro-F1 performance. Further analysis shows that errors concentrate in high-impact entities despite relatively stable action-type classification, indicating structural grounding deficiencies. These findings highlight the necessity of consequence-aware evaluation protocols for the responsible deployment of AI-assisted ATC systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。