用事故报告生成逼真工业危险场景图,提升安全训练数据质量
Scene Graph-Guided Generative AI Framework for Synthesizing and Evaluating Industrial Hazard Scenarios
- 基于OSHA报告提取结构化风险信息,构建物体级场景图
- 场景图引导扩散模型生成符合空间关系的危险场景图像
- 自研VQA评分比CLIP/BLIP更敏感,可有效评估生成图像真实度
训练视觉模型准确检测工作场所危险需要真实反映潜在事故场景的图像,但这类数据难以获取,因实际事故过程几乎无法捕捉。为此,本文提出一种基于场景图的生成式AI框架,利用历史美国职业安全与健康管理局(OSHA)事故报告合成逼真危险场景图像。通过GPT-4o分析OSHA文本,提取结构化危害推理,并转化为包含空间与上下文关系的对象级场景图,指导文本到图像扩散模型生成构图准确的危险场景。为评估生成数据的真实性和语义保真度,引入一种视觉问答(VQA)框架。在四种主流生成模型上,所提出的VQA图得分优于基于熵的CLIP和BLIP指标,证实其具备更高判别灵敏度。
原文摘要 · Abstract (English)
Training vision models to detect workplace hazards accurately requires realistic images of unsafe conditions that could lead to accidents. However, acquiring such datasets is difficult because capturing accident-triggering scenarios as they occur is nearly impossible. To overcome this limitation, this study presents a novel scene graph-guided generative AI framework that synthesizes photorealistic images of hazardous scenarios grounded in historical Occupational Safety and Health Administration (OSHA) accident reports. OSHA narratives are analyzed using GPT-4o to extract structured hazard reasoning, which is converted into object-level scene graphs capturing spatial and contextual relationships essential for understanding risk. These graphs guide a text-to-image diffusion model to generate compositionally accurate hazard scenes. To evaluate the realism and semantic fidelity of the generated data, a visual question answering (VQA) framework is introduced. Across four state-of-the-art generative models, the proposed VQA Graph Score outperforms CLIP and BLIP metrics based on entropy-based validation, confirming its higher discriminative sensitivity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。