arXiv:2506.19342cs.LGcs.AI2025-06被引 1

通过文本对齐发现车祸数据中酒精判断错误,提升报告准确性。

Unlocking Insights Addressing Alcohol Inference Mismatch through Database-Narrative Alignment

  • 用BERT模型对撞车记录的文本与数据库比对,识别酒精误判。
  • 分析37万条数据发现24.03%事故存在酒精判断错误。
  • 夜间、酒驾致死事故错判少,老人或车型不明事故更易出错。

道路交通事故是全球主要致死原因,亟需准确数据以支持预防策略和政策制定。本研究针对酒精判断不一致(AIM)问题,采用数据库与文本信息对齐方法识别事故数据中的错误。基于BERT模型分析爱荷华州2016至2022年共371,062条事故记录,发现2,767起存在酒精判断不一致,整体占比达24.03%。结合Probit Logit模型分析显示,酒驾致死事故及夜间事故的错判率较低,而涉及未知车辆类型或老年驾驶员的事故更易出现不一致。研究还通过地理空间聚类识别出需加强培训的重点区域。结果表明,应加强针对性培训与数据管理,提升事故报告准确性,支撑科学决策。

原文摘要 · Abstract (English)

Road traffic crashes are a significant global cause of fatalities, emphasizing the urgent need for accurate crash data to enhance prevention strategies and inform policy development. This study addresses the challenge of alcohol inference mismatch (AIM) by employing database narrative alignment to identify AIM in crash data. A framework was developed to improve data quality in crash management systems and reduce the percentage of AIM crashes. Utilizing the BERT model, the analysis of 371,062 crash records from Iowa (2016-2022) revealed 2,767 AIM incidents, resulting in an overall AIM percentage of 24.03%. Statistical tools, including the Probit Logit model, were used to explore the crash characteristics affecting AIM patterns. The findings indicate that alcohol-related fatal crashes and nighttime incidents have a lower percentage of the mismatch, while crashes involving unknown vehicle types and older drivers are more susceptible to mismatch. The geospatial cluster as part of this study can identify the regions which have an increased need for education and training. These insights highlight the necessity for targeted training programs and data management teams to improve the accuracy of crash reporting and support evidence-based policymaking.

交通安全数据质量自然语言处理事故分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。