arXiv:2603.11987cs.AI2026-03被引 2

构建实验室安全推理基准,评估AI在高危实验中的风险识别能力。

LABSHIELD: A Multimodal Benchmark for Safety-Critical Reasoning and Planning in Scientific Laboratories

  • 基于OSHA和GHS标准,设计164项多任务安全评估体系。
  • 20个商用模型平均在专业场景下安全表现下降32.0%。
  • 适合研究自主实验机器人安全决策的学者使用。

人工智能正加速科学自动化,多模态大语言模型(MLLM)代理已从实验室助手演变为自主实验操作者。这一转变对实验室环境提出严苛的安全要求,因脆弱玻璃器皿、危险物质及高精度设备的存在,规划错误或风险误判可能导致不可逆后果。然而,实体智能体在高风险场景中的安全意识与决策可靠性仍缺乏明确界定与评估。为此,我们提出LABSHIELD,一个基于多视角的真实世界基准,用于评估MLLM在危害识别与安全关键推理方面的能力。该基准依据美国职业安全与健康管理局(OSHA)标准与全球化学品统一分类和标签制度(GHS),建立涵盖164项操作任务的严谨安全分类体系,任务涉及多样化的操作复杂度与风险特征。我们在双轨评估框架下测试了20个专有模型、9个开源模型及3个实体模型。结果揭示,通用领域MCQ准确率与半开放问答安全性能之间存在系统性差距,模型在专业实验室场景中平均性能下降32.0%,尤其在危害解读与安全感知规划方面表现薄弱。这些发现凸显了构建以安全为核心的推理框架的紧迫性,以保障实体实验室环境中自主科学实验的可靠性。完整数据集即将发布。

原文摘要 · Abstract (English)

Artificial intelligence is increasingly catalyzing scientific automation, with multimodal large language model (MLLM) agents evolving from lab assistants into self-driving lab operators. This transition imposes stringent safety requirements on laboratory environments, where fragile glassware, hazardous substances, and high-precision laboratory equipment render planning errors or misinterpreted risks potentially irreversible. However, the safety awareness and decision-making reliability of embodied agents in such high-stakes settings remain insufficiently defined and evaluated. To bridge this gap, we introduce LABSHIELD, a realistic multi-view benchmark designed to assess MLLMs in hazard identification and safety-critical reasoning. Grounded in U.S. Occupational Safety and Health Administration (OSHA) standards and the Globally Harmonized System (GHS), LABSHIELD establishes a rigorous safety taxonomy spanning 164 operational tasks with diverse manipulation complexities and risk profiles. We evaluate 20 proprietary models, 9 open-source models, and 3 embodied models under a dual-track evaluation framework. Our results reveal a systematic gap between general-domain MCQ accuracy and Semi-open QA safety performance, with models exhibiting an average drop of 32.0% in professional laboratory scenarios, particularly in hazard interpretation and safety-aware planning. These findings underscore the urgent necessity for safety-centric reasoning frameworks to ensure reliable autonomous scientific experimentation in embodied laboratory contexts. The full dataset will be released soon.

实验室自动化多模态模型安全评估AI推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。