arXiv:2604.08987cs.AI2026-04中稿 · the 2026 IEEE Inte…被引 2

测试大模型在飞行安全约束下的物理推理能力,发现其可控性强但精度不足。

PilotBench: A Benchmark for General Aviation Agents with Safety Constraints

  • 构建包含708条真实飞行轨迹的基准,同步采集34通道遥测数据
  • 传统模型精度高(MAE 7.01),大模型指令遵循率达86%-89%但精度下降至11-14
  • 高负荷阶段性能骤降,提示需结合符号推理与数值预测的混合架构

随着大语言模型向具身智能代理发展,一个核心问题浮现:基于文本训练的模型能否在遵守安全约束的前提下可靠地进行复杂物理推理?为此,我们提出PilotBench,一个评估大语言模型在安全关键型飞行轨迹与姿态预测任务中表现的基准。该基准基于708条真实通用航空飞行轨迹,涵盖九个操作上不同的飞行阶段,并同步记录34通道遥测数据。通过对比大语言模型与传统预测器,系统检验了语义理解与物理驱动预测的交叉能力。我们引入Pilot-Score复合指标,综合平衡60%回归准确率与40%指令遵循及安全合规性。41种模型的对比分析揭示了‘精度-可控性二分法’:传统预测器实现7.01的低MAE,但缺乏语义推理能力;大语言模型在指令遵循率86–89%的情况下,精度下降至11–14。分阶段分析进一步暴露‘动态复杂性差距’——大模型在爬升、进近等高负荷阶段性能显著下滑,表明其隐式物理模型脆弱。这些实证发现推动了融合大模型符号推理与专用预测器数值精度的混合架构发展。PilotBench为安全约束领域具身智能的演进提供了严谨基础。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) advance toward embodied AI agents operating in physical environments, a fundamental question emerges: can models trained on text corpora reliably reason about complex physics while adhering to safety constraints? We address this through PilotBench, a benchmark evaluating LLMs on safety-critical flight trajectory and attitude prediction. Built from 708 real-world general aviation trajectories spanning nine operationally distinct flight phases with synchronized 34-channel telemetry, PilotBench systematically probes the intersection of semantic understanding and physics-governed prediction through comparative analysis of LLMs and traditional forecasters. We introduce Pilot-Score, a composite metric balancing 60% regression accuracy with 40% instruction adherence and safety compliance. Comparative evaluation across 41 models uncovers a Precision-Controllability Dichotomy: traditional forecasters achieve superior MAE of 7.01 but lack semantic reasoning capabilities, while LLMs gain controllability with 86--89% instruction-following at the cost of 11--14 MAE precision. Phase-stratified analysis further exposes a Dynamic Complexity Gap-LLM performance degrades sharply in high-workload phases such as Climb and Approach, suggesting brittle implicit physics models. These empirical discoveries motivate hybrid architectures combining LLMs' symbolic reasoning with specialized forecasters' numerical precision. PilotBench provides a rigorous foundation for advancing embodied AI in safety-constrained domains.

具身智能飞行预测安全约束大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。