arXiv:2412.02094cs.LGcs.CY2024-12被引 2

针对大型车辆作业区事故数据不平衡问题,提出高效预测严重事故的建模策略。

Crash Severity Risk Modeling Strategies under Data Imbalance

  • 用互信息筛选关键特征,提升高严重性事故预测能力
  • 近邻欠采样结合特定模型可使高严重性事故召回率达最高
  • 不同安全目标下推荐匹配模型与数据处理方法

本研究针对南卡罗来纳州2014至2018年大型车辆作业区事故数据中低严重性(LS)事故数量是高严重性(HS)事故四倍的数据不平衡问题,评估多种统计、机器学习与深度学习模型在不同特征选择与数据平衡技术下的事故严重程度预测性能。结果表明,由于类别不平衡与特征重叠,高严重性事故预测准确率较低。判别式互信息(DMI)能有效提取用于预测高严重性事故的特征集,无需数据平衡,尤其配合梯度提升模型(如CatBoost、XGBoost、LightGBM)及神经网络(NeuralNetTorch)表现最优。近邻欠采样(NearMiss-1)结合DMI特征可最大化高严重性事故召回率,适用于高严重性预测任务。而随机欠采样、高严重性类别加权与随机过采样则实现更均衡的性能表现,即低与高严重性指标间的合理权衡,尤其在融合特征集或无特征选择的NeuralNetTorch、NeuralNetFastAI、CatBoost、LightGBM与贝叶斯混合逻辑回归(BML)模型上效果显著。研究为安全分析人员提供了依据具体安全目标选择模型、特征筛选与数据平衡方法的指导,为提升作业区事故严重性预测能力提供可靠基础。

原文摘要 · Abstract (English)

This study investigates crash severity risk modeling strategies for work zones involving large vehicles (i.e., trucks, buses, and vans) under crash data imbalance between low-severity (LS) and high-severity (HS) crashes. We utilized crash data involving large vehicles in South Carolina work zones from 2014 to 2018, which included four times more LS crashes than HS crashes. The objective of this study is to evaluate the crash severity prediction performance of various statistical, machine learning, and deep learning models under different feature selection and data balancing techniques. Findings highlight a disparity in LS and HS predictions, with lower accuracy for HS crashes due to class imbalance and feature overlap. Discriminative Mutual Information (DMI) yields the most effective feature set for predicting HS crashes without requiring data balancing, particularly when paired with gradient boosting models and deep neural networks such as CatBoost, NeuralNetTorch, XGBoost, and LightGBM. Data balancing techniques such as NearMiss-1 maximize HS recall when combined with DMI-selected features and certain models such as LightGBM, making them well-suited for HS crash prediction. Conversely, RandomUnderSampler, HS Class Weighting, and RandomOverSampler achieve more balanced performance, which is defined as an equitable trade-off between LS and HS metrics, especially when applied to NeuralNetTorch, NeuralNetFastAI, CatBoost, LightGBM, and Bayesian Mixed Logit (BML) using merged feature sets or models without feature selection. The insights from this study offer safety analysts guidance on selecting models, feature selection, and data balancing techniques aligned with specific safety goals, providing a robust foundation for enhancing work-zone crash severity prediction.

事故预测数据不平衡机器学习交通安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。