arXiv:2506.12092cs.CLcs.LG2025-06被引 1

用NLP分析事故文本,提升城市交通安全分类准确率

Enhancing Traffic Accident Classifications: Application of NLP Methods for City Safety

  • 用主题建模和少样本学习分析事故文本标签一致性
  • 文本描述比结构化数据更关键,模型准确率高
  • 适合交通管理与政策制定者参考

全面理解交通事故对提升城市安全和制定政策至关重要。本研究以慕尼黑的交通事故数据为基础,分析不同事故类型的模式与特征。数据包含结构化特征(如位置、时间、天气)及非结构化文本描述,每起事故被划分为七个预定义类别。通过应用NLP方法(包括主题建模和少样本学习),我们发现标签存在不一致问题,揭示了分类过程中的潜在模糊性,并推动了更精确的预测方法。基于此,我们构建了一个分类模型,在事故类别识别上表现优异。结果表明,文本描述是最重要的信息来源,而结构化数据仅带来微弱改进。这凸显了自由文本在事故分析中的关键作用,并展示了基于Transformer的模型在提升分类可靠性方面的潜力。

原文摘要 · Abstract (English)

A comprehensive understanding of traffic accidents is essential for improving city safety and informing policy decisions. In this study, we analyze traffic incidents in Munich to identify patterns and characteristics that distinguish different types of accidents. The dataset consists of both structured tabular features, such as location, time, and weather conditions, as well as unstructured free-text descriptions detailing the circumstances of each accident. Each incident is categorized into one of seven predefined classes. To assess the reliability of these labels, we apply NLP methods, including topic modeling and few-shot learning, which reveal inconsistencies in the labeling process. These findings highlight potential ambiguities in accident classification and motivate a refined predictive approach. Building on these insights, we develop a classification model that achieves high accuracy in assigning accidents to their respective categories. Our results demonstrate that textual descriptions contain the most informative features for classification, while the inclusion of tabular data provides only marginal improvements. These findings emphasize the critical role of free-text data in accident analysis and highlight the potential of transformer-based models in improving classification reliability.

交通分析NLP应用事故分类文本挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。