融合图像与事故数据,用AI评估铁路交叉口安全等级。
Multi-modal Rail Crossing Safety Analysis

- 结合视觉图像与历史事故数据,构建多模态安全评估模型。
- 高危/低危交叉口识别F1达0.757,安全评分与FRA标准相关性0.492。
- 结果符合专家判断,适合交通安全部门和智能运维场景使用。
给定一个或多个铁路交叉口的图像,能否利用视觉线索可靠地估算其安全性?若引入该交叉口的官方事故报告等结构化历史数据,能否提升评估能力?本文探索上述问题,旨在构建一个可处理多模态数据的AI系统,为铁路交叉口提供与专家意见及联邦铁路管理局(FRA)安全评分一致的安全评估与打分。为此,我们提出一个概念验证流程,涵盖从数据准备到不同学习范式的全链条挑战。实验表明,所提系统在路由微调的紧凑视觉语言模型(VLM)管道下,实现高危/低危交叉口识别的宏平均F1为0.757,对FRA安全评分的均方根误差(RMSE)为0.078,相关系数达0.492,且生成结果与领域专家评估一致。
原文摘要 · Abstract (English)
Given one or more images of a railway crossing, can we leverage visual cues that allow us to robustly estimate how safe it is? Can we improve our ability to do so by introducing structured data (such as official accident reports) about the accident history of that crossing into our models? In this work, we explore how to best answer those questions towards building an AI system that can ingest multi-modal data for railway crossings and provide safety assessment and scores that align with expert opinion and with safety scoring used by the Federal Railroad Administration (FRA). To that end, we propose a proof-of-concept pipeline that delivers on that goal, while at the same time exploring and tackling a number of critical research challenges that pertain to different parts of the pipeline, from data preparation to different learning paradigms that can allow us to realize such a system. Indicatively, our proposed system identifies HIGH-RISK and LOW-RISK crossings with a macro F1 score of 0.757 and estimates FRA-based safety scores with an RMSE of 0.078 and correlation of 0.492 using a routed fine-tuned compact VLM pipeline, while producing qualitative results that align with domain-expert assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。