用定制大模型分析车祸数据,让预测更准、原因更清楚。
Towards Reliable and Interpretable Traffic Crash Pattern Prediction and Safety Interventions Using Customized Large Language Models
- 把车祸文本、图像等数据转为文本,用大模型做推理预测
- 相比基线模型F1提升42%,能精准识别严重事故主因
- 可解释每条预测依据,适合交通政策制定者和安全研究者
预测交通事故对理解事故分布及其影响因素至关重要,有助于设计主动交通安全干预措施。然而,现有方法难以解析数值特征、文本报告、事故图像、环境条件和驾驶行为记录等多种来源数据之间的复杂关系,往往无法捕捉其中丰富的语义信息和内在关联,限制了对关键风险因素的识别。本研究提出TrafficSafe框架,通过适配大语言模型(LLM),将车祸预测与特征归因转化为基于文本的推理任务。构建了一个包含58,903条真实事故报告的多模态数据集,并将其文本化为TrafficSafe Event Dataset。通过对该数据集定制化微调,TrafficSafe LLM在F1-score上相较基线平均提升42%。为解释预测结果并挖掘贡献因素,引入TrafficSafe Attribution——一种句级特征归因框架,支持条件性风险分析。结果显示,酒驾是严重事故的主要因素,攻击性和受损行为对严重事故的贡献几乎为其他驾驶行为的两倍。此外,该框架在训练中识别出关键特征,指导迭代式数据收集以持续提升性能。TrafficSafe为交通安全管理提供了可解释、可行动的智能解决方案。
原文摘要 · Abstract (English)
Predicting crash events is crucial for understanding crash distributions and their contributing factors, thereby enabling the design of proactive traffic safety policy interventions. However, existing methods struggle to interpret the complex interplay among various sources of traffic crash data, including numeric characteristics, textual reports, crash imagery, environmental conditions, and driver behavior records. As a result, they often fail to capture the rich semantic information and intricate interrelationships embedded in these diverse data sources, limiting their ability to identify critical crash risk factors. In this research, we propose TrafficSafe, a framework that adapts LLMs to reframe crash prediction and feature attribution as text-based reasoning. A multi-modal crash dataset including 58,903 real-world reports together with belonged infrastructure, environmental, driver, and vehicle information is collected and textualized into TrafficSafe Event Dataset. By customizing and fine-tuning LLMs on this dataset, the TrafficSafe LLM achieves a 42% average improvement in F1-score over baselines. To interpret these predictions and uncover contributing factors, we introduce TrafficSafe Attribution, a sentence-level feature attribution framework enabling conditional risk analysis. Findings show that alcohol-impaired driving is the leading factor in severe crashes, with aggressive and impairment-related behaviors having nearly twice the contribution for severe crashes compared to other driver behaviors. Furthermore, TrafficSafe Attribution highlights pivotal features during model training, guiding strategic crash data collection for iterative performance improvements. The proposed TrafficSafe offers a transformative leap in traffic safety research, providing a blueprint for translating advanced AI technologies into responsible, actionable, and life-saving outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。