测试大模型在实验室安全任务中的表现,发现准确率普遍不足70%。
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
- 构建涵盖3128个任务的实验室安全评测基准
- 19个模型在危险识别上最高仅达69.5%准确率
- 闭源模型在开放推理中无明显优势,需更严格评估
人工智能正重塑科研方式,但其在实验室环境中的应用带来严峻安全挑战。大语言模型(LLMs)和视觉语言模型(VLMs)虽能协助实验设计与操作指导,但其‘理解幻觉’可能导致研究人员过度信赖不安全输出。本文提出LabSafety Bench,一个全面的评测基准,涵盖765道多选题与404个真实实验室场景,共3128个开放任务,用于评估模型在危险识别、风险评估与后果预测方面的能力。对19个先进LLMs和VLMs的评估显示,无一模型在危险识别任务上准确率超过70%。尽管闭源模型在结构化测试中表现较好,但在开放推理任务中未展现显著优势。结果表明,在真实实验室部署AI前,亟需建立专用的安全评估框架。
原文摘要 · Abstract (English)
Artificial Intelligence (AI) is revolutionizing scientific research, yet its growing integration into laboratory environments presents critical safety challenges. Large language models (LLMs) and vision language models (VLMs) now assist in experiment design and procedural guidance, yet their "illusion of understanding" may lead researchers to overtrust unsafe outputs. Here we show that current models remain far from meeting the reliability needed for safe laboratory operation. We introduce LabSafety Bench, a comprehensive benchmark that evaluates models on hazard identification, risk assessment, and consequence prediction across 765 multiple-choice questions and 404 realistic lab scenarios, encompassing 3,128 open-ended tasks. Evaluations on 19 advanced LLMs and VLMs show that no model evaluated on hazard identification surpasses 70% accuracy. While proprietary models perform well on structured assessments, they do not show a clear advantage in open-ended reasoning. These results underscore the urgent need for specialized safety evaluation frameworks before deploying AI systems in real laboratory settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。