用AI生成高速公路施工危险的合成图像和动态序列,解决安全培训素材稀缺问题。
Generative AI for Visualizing Highway Construction Hazards Through Synthetic Images and Temporal Sequences

- 基于事故报告文本生成单图或四阶段动态图像。
- 单图教育可用性达81.1%,时序图像对齐度3.94/5。
- 首次实现从伤亡报告生成时序危险可视化,适合安全培训与研究者使用。
高速公路施工人员面临严重受伤或死亡的高风险。基于图像的安全培训材料对提升培训效果至关重要,但因伦理与物流限制,现有资料稀缺。本研究开发并评估了一种生成式AI方法,从OSHA严重伤害报告文本中生成高速公路施工危险的合成可视化内容。提出两种模式:单次生成(每起事件生成一张图像)和时序生成(生成四阶段动态序列)。选取75条事故记录,共生成750张图像,通过CLIP语义检索与专家评估(教育实用性、保真度、一致性等维度)进行验证。单图模式在教育接受度上达到81.1%,保真度与一致性得分分别为4.14/5和4.07/5;时序序列接受度为60.9%,一致性得分3.94/5,但保真度较低(3.51/5)。CLIP检索结果显示两种模式均具备显著的语义检索能力。本研究是首批利用自回归图像生成模型从真实伤亡报告中生成施工危险视觉内容的工作,并首次实现时序化危险场景生成,同时构建了多维度评估框架,可推广至其他领域。该方法使安全培训无需实拍危险场景即可结合叙事与视觉教学,具有广泛应用潜力。
原文摘要 · Abstract (English)
Highway construction workers face a high risk of serious injury or death. Image-based training materials depicting hazardous scenarios are essential for engaging safety instruction but remain scarce due to ethical and logistical barriers. This study develops and evaluates a generative AI methodology for producing synthetic visualizations of highway construction hazards from OSHA Severe Injury Report narratives. Two modes were developed: a single-pass approach yielding one image per incident, and a temporal approach producing a four-stage sequence. A sample of 75 incident records yielded 750 images, evaluated using CLIP-based semantic retrieval and expert assessment across dimensions such as educational utility, fidelity, and alignment. Single-pass images achieved 81.1% educational acceptability with fidelity and alignment scores of 4.14/5 and 4.07/5, respectively, while temporal sequences achieved 60.9% acceptability with comparable alignment (3.94/5) but lower fidelity (3.51/5). CLIP-based retrieval revealed that both modes produce images with statistically significant retrieval capabilities. This is among the first studies to leverage modern autoregressive image generation models for visualizing construction hazards from reported severe injuries and to generate temporally sequenced hazard imagery, and a new multi-dimensional evaluation framework was developed to support future research in this domain. The work enables safety trainers to pair narrative storytelling with visual learning material without photographing real-world hazards, and the framework could be applied to datasets across diverse domains, enabling synthetic image generation tailored to new application areas.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。