用真人参与的监管级测试框架评估空管AI,更真实可靠。
Human-in-the-Loop Testing of AI Agents for Air Traffic Control with a Regulated Assessment Framework
- 引入监管认证模拟器和真人教官参与评估
- 实现与真实空管环境对齐的性能测量
- 适合空管AI研发与人机协同研究者参考
我们提出一种严格的、基于真人参与的评估框架,用于衡量AI代理在空中交通管制任务中的表现。该框架依托于经监管机构认证的仿真训练课程,该课程同样用于培训真实空管员。通过采用法定合规的评估流程并引入专家人类教官参与,该框架实现了对AI性能更为真实、领域准确的度量。本研究解决了现有文献中一个关键缺陷:学术界对空管任务的简化建模与实际运行环境复杂性之间的严重脱节。同时,该工作为未来高效的人机协同范式奠定了基础,使机器性能与人类评估目标保持一致。
原文摘要 · Abstract (English)
We present a rigorous, human-in-the-loop evaluation framework for assessing the performance of AI agents on the task of Air Traffic Control, grounded in a regulator-certified simulator-based curriculum used for training and testing real-world trainee controllers. By leveraging legally regulated assessments and involving expert human instructors in the evaluation process, our framework enables a more authentic and domain-accurate measurement of AI performance. This work addresses a critical gap in the existing literature: the frequent misalignment between academic representations of Air Traffic Control and the complexities of the actual operational environment. It also lays the foundations for effective future human-machine teaming paradigms by aligning machine performance with human assessment targets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。