arXiv:2511.10027cs.AI2025-11

测试大模型在化学品应急响应中的实用能力,发现有潜力但需人工把关。

ChEmREF: Evaluating Language Model Readiness for Chemical Emergency Response

  • 构建新基准ChEmREF,涵盖1035种化学品的三类任务
  • 最高准确率68%用于化学式与名称互译,63.9%通过安全考试
  • 适合应急决策系统研发者和化学品安全领域研究者参考

应对危险品事故的应急人员面临关键、紧迫的决策挑战,需手动查阅大量化学品指南。本文探讨当前语言模型能否快速可靠地理解关键信息、识别危害并提供建议。我们提出了化学应急响应评估框架(ChEmREF),包含来自《应急响应指南》和PubChem数据库的1,035种危险品的三个任务:(1) 化学表示形式在结构化与非结构化间的转换(如将C2H6O转为乙醇),(2) 应急响应生成(如推荐撤离距离),(3) 化学品安全与认证考试中的知识问答。最佳模型在非结构化化学表示转换上达到68.0%的精确匹配,在事件响应建议上获得LLM Judge评分52.7%,在化学品考试中多选题准确率达63.9%。结果表明,尽管语言模型在多项任务中展现潜力,但仍需谨慎的人工监督以应对当前局限。

原文摘要 · Abstract (English)

Emergency responders managing hazardous material HAZMAT incidents face critical, time-sensitive decisions, manually navigating extensive chemical guidelines. We investigate whether today's language models can assist responders by rapidly and reliably understanding critical information, identifying hazards, and providing recommendations. We introduce the Chemical Emergency Response Evaluation Framework (ChEmREF), a new benchmark comprising questions on 1,035 HAZMAT chemicals from the Emergency Response Guidebook and the PubChem Database. ChEmREF is organized into three tasks: (1) translation of chemical representation between structured and unstructured forms (e.g., converting C2H6O to ethanol), (2) emergency response generation (e.g., recommending appropriate evacuation distances) and (3) domain knowledge question answering from chemical safety and certification exams. Our best evaluated models received an exact match of 68.0% on unstructured HAZMAT chemical representation translation, a LLM Judge score of 52.7% on incident response recommendations, and a multiple-choice accuracy of 63.9% on HAMZAT examinations. These findings suggest that while language models show potential to assist emergency responders in various tasks, they require careful human oversight due to their current limitations.

应急响应语言模型化学品安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。