arXiv:2607.20476cs.AIcs.CL2026-07

五款大模型在多传感器安全预警中集体失灵,却能精准识别单传感器越限。

Benchmarking Large Language Models on Multi-Sensor Physical Hazard Assessment

论文配图:Benchmarking Large Language Models on Multi-Sensor Physical Hazard Assessment
图 1 · 摘自论文原文
  • 测试5个大模型在多传感器同时轻微超限场景下的预警能力
  • 多传感器联合场景下预警准确率近乎为零,单传感器越限时准确率达97.5%以上
  • 文本格式对结果无显著影响,纯文本表述更优,适合安全监控系统部署

我们开展了一项实证基准测试,评估五款大语言模型在多传感器物理危险数据评估中的表现。在温度0.0下,通过1800次API调用测试了60个场景,涵盖三类任务:多传感器联合评估、响应比例性以及模式消歧。结果显示,当多个传感器同时轻微超限但未突破各自安全阈值时,所有模型均未发出任何预警信号,而单传感器越限时准确率接近完美(类别B Q1:0.975-1.000)。在类别A的多传感器场景中,各模型得分极低(Q2:0.000-0.208;Q3:0.000-0.592),且结构化表格格式未带来明显优势,纯文本表述反而使ChatGPT-4o表现更佳(p = 0.001)。这些发现对实际部署于物理安全监测系统的大模型应用具有直接指导意义。

原文摘要 · Abstract (English)

We present an empirical benchmark evaluating how five large language models assess multisensor physical hazard data. Testing 60 scenarios across three categories - multi-sensor joint assessment, response proportionality, and pattern disambiguation - with 1,800 API calls at temperature 0.0, we find that all tested models consistently produced no precautionary warning signal across the tested scenarios where multiple sensors are simultaneously elevated below their individual safety limits, while achieving near-perfect accuracy on single-sensor threshold violations. All five models (ChatGPT-4o, Gemini 2.5 Flash, DeepSeek, Kimi, Llama 3.1 8B) score near zero on Category A multi-sensor scenarios (Q2: 0.000-0.208; Q3: 0.000-0.592) compared to strong performance on single-sensor scenarios (Category B Q1: 0.975-1.000). Structured tabular formatting shows no consistent advantage over plain prose; ChatGPT-4o performs significantly better under prose (p = 0.001). These findings have direct implications for practitioners deploying the tested models in physical safety monitoring systems.

大模型评测安全预警多传感器可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。