构建航空操作知识评测基准,检验大模型在真实场景下的可靠性。
Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge
- 基于国际标准与机场运营材料设计300道多选题
- 最强模型准确率仅82.7%,距专家水平95%仍有差距
- 专为航空领域安全推理设计,适合监管与落地评估
大语言模型在航空业务中的应用日益广泛,涵盖文档生成、培训内容制作及面向客户的助手。然而通用基准无法衡量模型在航空特定操作知识上的安全与正确推理能力,而该领域的高风险与强监管特性使得这一差距尤为严重。本文提出 Pre-Flight,一个开源基准,包含300道源自国际标准和机场地面运营资料的多选题,覆盖国际机场地面运行、ICAO与美国FAA法规、航空通用知识及复杂运行场景。题目由空管、地面运营和商业飞行从业者共同撰写与审核。我们采用 Inspect 评估框架,对一系列主流商业与开源模型进行评估,按标准多选协议评分,并持续更新排行榜。以航空专业人士小样本测试获得的约95%为非正式专家基准,即使最新发布的2026年最强模型也仅达82.7%,相比2025年初约75%的水平改善有限。显著且持续存在的性能差距表明,当前模型尚未达到专家级可靠性。我们公开数据集、评估工具与结果,基准已集成至 inspect_evals 社区评估包中。我们认为,此类领域专属评估是生成式AI在非安全关键航空操作中负责任部署的必要前提。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly proposed for aviation business operations, from documentation and training generation to customer facing assistants. General purpose benchmarks do not measure whether a model reasons safely and correctly about aviation specific operational knowledge, and the high stakes, regulated nature of the domain makes that gap consequential. We present Pre-Flight, an open source benchmark of 300 multiple choice questions drawn from international standards and airport ground operations material, covering international airport ground operations, ICAO and US FAA regulations, aviation general knowledge and complex operational scenarios. Questions were authored and reviewed by practitioners with experience in air traffic management, ground operations and commercial flying. We evaluate a range of contemporary commercial and open weight models using the Inspect evaluation framework, scoring by accuracy under a standard multiple choice protocol, and we maintain the leaderboard on a rolling basis as new models are released. Against an informal expert reference of around 95%, obtained from a low sample quiz of aviation professionals at a conference, even the strongest model evaluated (released in 2026) reaches 82.7%, having improved only gradually from roughly 75% in early 2025. A substantial and persistent gap below expert level reliability therefore remains. We release the dataset, the evaluation harness and the results, and the benchmark is available within the community evaluations package distributed with inspect_evals. We argue that domain specific evaluation of this kind is a necessary precondition for responsible deployment of generative AI in non safety critical aviation operations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。