arXiv:2608.24621cs.CL2026-08

语言模型在航空管制中看似准确,实则存在安全风险盲区。

Beyond Semantic Accuracy: Consequence-Aware Evaluation for Safety-Critical Language Understanding

论文配图:Beyond Semantic Accuracy: Consequence-Aware Evaluation for Safety-Critical Language Understanding
图 1 · 摘自论文原文
  • 用后果感知评估框架检验模型真实安全性
  • 8个模型在标准指标上表现良好,但安全评估结果差
  • 适合评估高风险场景下的AI可靠性

语言模型能否在安全关键任务中被信任?在这些场景中,语义指标表现优秀并不意味着实际操作可靠:读错高度、遗漏执行条件或混淆呼号,虽在标准F1指标上得分较高,却可能带来严重后果。本文以空中交通管制(ATC)为场景,该领域对错误容忍度近乎为零,采用后果感知评估方法检验语义分数是否虚高。研究基于航空标准构建可控诊断基准,并整合来自三个国家40名空中交通管制员的反馈。评估8个模型后发现,存在系统性语义-安全差距:即使模型在常规指标上表现可靠,其真实安全性能仍被显著高估。风险感知微调可缩小但无法消除该差距,表明后果感知评估是部署前不可或缺的补充,必须与传统NLP指标并行使用。

原文摘要 · Abstract (English)

Can language models be trusted in safety- critical operations? In such settings, strong per- formance on semantic metrics does not guaran- tee operational reliability: a misread altitude, a dropped execution condition, or a confused call- sign may score well under standard F1 yet carry sharply asymmetric operational consequences. We study this problem in air traffic control (ATC), where controller-pilot communication demands near-zero error tolerance, and use consequence-aware evaluation to test whether semantic scores misstate operational reliabil- ity. The framework is instantiated in a con- trolled diagnostic ATC benchmark grounded in aviation standards and feedback from 40 air traffic controllers across three countries. Evaluating 8 models, we uncover a system- atic semantic-safety gap: conventional scores give substantially higher performance estimates than consequence-aware evaluation, even for models that appear reliable under standard met- rics. Risk-aware fine-tuning narrows but does not close this gap, showing that consequence- aware evaluation is a necessary complement to standard NLP metrics before any real safety- critical deployment claim

安全评估航空管制语言模型后果感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。