用大模型分析施工事故报告,自动分类并推荐相关安全资料。
AISA: AI Safety Assistant Framework for Continuous Improvement of Highway Construction

- 用大模型对事故文本进行多类分类和质量评分
- 事故检索准确率超随机水平,尤其对具体施工活动有效
- 支持本地推理,适合实际工地持续改进安全计划
作业安全分析(JSA)与任务前规划可受益于历史事故记录,但这些数据常以非结构化叙述形式存在,难以在规划时快速查阅。本文提出一种基于大语言模型(LLM)的高速公路施工安全报告与规划框架,为未来智能体应用奠定基础,强调确定性与本地推理。第一目标是实现事故叙述的分类与质量评分,用于现有及未来报告;第二目标是评估相关历史事故、关联图像与可信行业文档的检索效果,融入日常安全计划。通过神经探针训练,对四个多分类和两个二分类的《职业伤害与疾病分类系统》(OIICS)字段进行分类,并生成综合质量评分,在超过15,000条叙述的测试集及100条人工标注样本上评估,对比多数投票的LLM集成模型。事故检索在不同嵌入模型间使用标准信息检索指标进行基准测试。OIICS分类在保留数据上达到75%的准确率,但两个二分类标签表现退化。质量评分在一个数据库中有效,但在外部分布的死亡事件上出现扭曲。事故检索显著高于随机水平,尤其在词汇差异明显的施工活动中表现最佳。文档问答中,开源解码器嵌入模型优于专有模型。整体上,该工作提供了一个基于本地推理与文本嵌入模型的新框架,强调将外部数据连接至JSA报告,适用于未来智能体应用。
原文摘要 · Abstract (English)
Job Safety Analysis (JSA) and pre-task planning can benefit from prior incident records, yet historical accident data is often stored as unstructured narratives that are difficult to consult at the point of planning. A novel framework centered on large language models (LLMs) for highway construction safety reporting and planning is proposed as a foundation for future agentic applications, prioritizing deterministic, local inferencing. The first aim is to enable classification and quality scoring of incident narratives for existing and future reporting purposes. The second is to evaluate retrieval of relevant historical accidents, related imagery, and trusted industry documents for incorporation into daily safety plans. Neural probes were trained to classify incidents along four multiclass and two binary Occupational Injury and Illness Classification System (OIICS) fields and to derive an overall quality score, evaluated on a test set of over 15,000 narratives and a held-out set of 100 author-labeled records, benchmarked against a majority-vote LLM ensemble. The retrieval of historical accidents, reference imagery, and industry documents was benchmarked across embedding models using standard information retrieval metrics. OIICS classification reached 75% held-out accuracy, though the two binary flags were degenerate. The quality score, while meaningful on one database, was distorted on out-of-distribution fatalities in the held-out dataset. Accident retrieval recovered relevant incidents far above chance, performing best on lexically distinct construction activities. On document question answering, an open-weight decoder embedding model surpassed proprietary models. Overall, this work provides a new framework rooted in local inferencing and text embedding models for future agentic applications, with emphasis on bridging external data to JSA reports.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。