构建工业安全多模态推理基准,连接场景与事故知识。
SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge

- 通过场景与报告双路径构建多模态安全问答数据
- 涵盖11万+场景类与1.3万+报告类问答对,覆盖因果分析与预防决策
- 揭示现有模型在多证据推理上的显著短板,适合安全智能研究者
工业安全理解不仅需检测人员、设备和防护装备,还需评估合规性、识别危险交互、解释潜在事故机制并提出预防措施。现有数据集多聚焦视觉感知或单一违规识别,缺乏证据支撑的推理监督。我们提出SafeSceneReason,一个连接工作场景与职业事故调查知识的多模态安全推理基准及配套训练语料库。该数据集采用两种互补构建路径:场景中心路径将标注的工作场景图像转化为可执行的安全场景图,并通过对象、关系与安全规则的程序执行生成确定性答案;报告中心路径从事故报告中提取图表与上下文证据,构建包含证据图、明确信息边界、多步推理路径与迭代验证的多模态问题。最终资源包含110,581个经验证的场景类问答对和13,114个优化的报告类问答对,覆盖感知、空间与量化推理、合规评估、证据整合、因果分析及以缓解为导向的决策制定。对代表性专有与开源视觉-语言模型的评估显示,其在比较性、技术性和多证据推理上存在显著性能差异与持续弱点,表明强大的通用视觉理解尚未能保证可靠的工业安全推理。
原文摘要 · Abstract (English)
Industrial-safety understanding requires more than detecting workers, equipment, and personal protective equipment. Models must also assess compliance, identify hazardous interactions, explain potential accident mechanisms, and recommend preventive actions. Existing safety datasets primarily focus on visual perception or isolated violation recognition and provide limited supervision for evidence-grounded reasoning. We introduce SafeSceneReason, a multimodal industrial-safety reasoning benchmark and companion training corpus that connects workplace scenes with knowledge from occupational accident investigations. SafeSceneReason combines two complementary data-construction pipelines. The scene-centric pipeline converts annotated workplace images into executable safety scene graphs and generates deterministic answers through program execution over objects, relations, and safety rules. The report-centric pipeline extracts figures and contextual evidence from accident reports and constructs multimodal questions using evidence graphs, explicit information boundaries, multi-step reasoning paths, and iterative verification. The resulting resource contains 110,581 verified scene-centric question--answer pairs and 13,114 refined report-centric question--answer pairs, covering perception, spatial and quantitative reasoning, compliance assessment, evidence synthesis, causal analysis, and mitigation-oriented decision making. Evaluation of representative proprietary and open-source vision--language models reveals substantial performance differences and persistent weaknesses in comparative, technical, and multi-evidence reasoning, demonstrating that strong general visual understanding does not yet guarantee reliable industrial-safety reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。