用分组特征变压器提升交通事故预测与因果分析精度
Feature Group Tabular Transformer: A Novel Approach to Traffic Crash Modeling and Causality Analysis
- 将多源数据分组为语义令牌,构建特征分组表格式变压器模型
- 在多种碰撞类型预测中优于随机森林等传统模型
- 揭示不同事故类型的主因,适合交通安全管理研究者
可靠的可解释交通事故建模对理解因果关系、提升道路安全至关重要。本研究提出一种新方法,利用融合天气数据、事故报告、高分辨率交通信息、路面几何及设施特征的综合数据集,预测碰撞类型。核心是开发特征分组表格式变压器(FGTT)模型,将异构数据组织为有意义的特征组,并以令牌形式表示。这些基于分组的令牌作为丰富语义单元,有效识别碰撞模式并解析因果机制。FGTT在多项指标上优于随机森林、XGBoost和CatBoost等主流树集成模型,且模型解释揭示了各类事故的关键影响因素,为不同碰撞类型的成因提供了新见解。
原文摘要 · Abstract (English)
Reliable and interpretable traffic crash modeling is essential for understanding causality and improving road safety. This study introduces a novel approach to predicting collision types by utilizing a comprehensive dataset fused from multiple sources, including weather data, crash reports, high-resolution traffic information, pavement geometry, and facility characteristics. Central to our approach is the development of a Feature Group Tabular Transformer (FGTT) model, which organizes disparate data into meaningful feature groups, represented as tokens. These group-based tokens serve as rich semantic components, enabling effective identification of collision patterns and interpretation of causal mechanisms. The FGTT model is benchmarked against widely used tree ensemble models, including Random Forest, XGBoost, and CatBoost, demonstrating superior predictive performance. Furthermore, model interpretation reveals key influential factors, providing fresh insights into the underlying causality of distinct crash types.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。