构建交互式飞行舱环境,评估大模型在真实操作中的安全执行能力。
AeroCopilotBench: A Two-Tier Benchmark for Evaluating LLM Agents as Aviation Copilots in an Interactive Virtual Cockpit Environment

- 设计双层评测框架,分知识测试与动态任务执行
- 12个模型中最高任务成功率72.6%,静态知识不保证实操能力
- 揭示流程不完整、状态反馈缺失等关键失败模式
大型语言模型(LLM)代理可能辅助机组进行复杂决策与任务执行,但现有航空评估多基于静态知识,无法系统检验交互环境中程序执行与安全合规性。本文提出AeroCopilot操作环境(ACOE),一个可复现的交互式虚拟驾驶舱测试环境,以及AeroCopilotBench双层航空代理评测基准。一级评测使用1,200道选择题评估航空知识,二级评测包含73项来自飞机操作手册(POH)的紧急与异常任务,嵌入ACOE中。ACOE将自然语言规程转化为可执行的状态转移、终态目标条件与硬性安全约束,使模型可通过标准工具接口理解驾驶舱状态、诊断故障并操控机载系统。我们建立安全门控评估框架:仅当所有任务目标达成且未违反任何硬性安全约束时,轨迹才算成功,同时分别度量安全进展与轨迹安全性。12个模型中,最高二级成功率72.6%,静态知识表现与程序执行能力无一致关联。对3个代表性模型451次失败案例分析发现,常见问题包括流程不完整、未利用状态反馈、长周期执行管理不足。这些结果推动状态感知的代理协同、任务完成与轨迹安全联合评估,以及重复回归测试。ACOE与AeroCopilotBench为航空代理的知识应用、交互执行与运行安全提供了可复现的测试基础。
原文摘要 · Abstract (English)
Large language model (LLM) agents may assist flight crews with complex decisions and task execution, but existing aviation evaluations centered on static knowledge do not support systematic testing of procedural execution and safety compliance in interactive environments. This paper presents the AeroCopilot Operational Environment (ACOE), a reproducible interactive virtual-cockpit test environment, and AeroCopilotBench, a two-tier aviation agent evaluation benchmark. Tier-1 evaluates aviation knowledge using 1,200 multiple-choice questions, while Tier-2 comprises 73 emergency and abnormal tasks derived from the manufacturers' Pilot's Operating Handbooks (POHs) and instantiated in ACOE. ACOE converts natural-language procedures into executable state transitions, final-state goal conditions, and hard safety constraints, enabling models to interpret cockpit state, diagnose faults, and operate aircraft systems through standardized tool interfaces. We establish a safety-gated evaluation framework in which a trajectory succeeds only when all task goals are achieved without violating any hard safety constraint, while safe goal progress and trajectory safety are measured separately. Across 12 models, the highest Tier-2 success rate is 72.6%, while static knowledge performance does not consistently translate into procedural execution. Analysis of 451 failed episodes from 3 representative models identifies recurring failures in procedural completeness, use of state feedback, and long-horizon execution management. These findings motivate state-aware agent orchestration, joint assessment of task completion and trajectory safety, and repeated regression testing. ACOE and AeroCopilotBench provide a reproducible foundation for testing knowledge application, interactive execution, and operational safety in aviation agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。