用警察事故描述测试大模型编码准确率,发现效果参差不齐。
Benchmarking Frontier Large Language Models Against Official Crash Database Coding Using Police Crash Narratives
- 用统一零样本提示让六款前沿大模型从事故叙述中提取属性
- 最高准确率仅达78.6%,光照条件编码最差
- 比不过简单规则基线,建议按属性逐项评估
警察事故叙述包含可补充结构化数据库的信息,但人工审核成本高,且大语言模型(LLMs)在还原官方事故编码方面的表现尚不明确。本研究通过对比阿肯色州致命事故数据库中的5,889条结构化记录与5,587条事故叙述,共匹配4,194起事故,对六款前沿大模型进行了基准测试。使用相同零样本提示,评估模型在事故方式、非机动车关系、交叉口类型、施工区关系、路面状况和光照条件等六项属性上的编码能力。评估指标包括一致性、宏平均F1分数、Cohen's kappa、覆盖率、选择性一致性和与始终多数、始终未知及关键词规则基线的比较。重复测量分析与广义估计方程模型用于评估模型与属性间的差异。GPT-5.5 High在所有模型中达成最高一致性,但始终多数基线产生更高原始一致性,关键词规则基线在宏平均F1分数和Cohen's kappa上与最佳模型相当。一致性在非机动车关系和事故方式上最高,光照条件、路面状况和施工区关系最低。属性间差异大于模型间差异。研究为基于大模型的事故编码评估提供了基准,并表明部署应基于属性进行透明基线比对与人工审查。
原文摘要 · Abstract (English)
Police crash narratives contain information that may supplement structured crash databases, but manual review is labor-intensive and it remains unclear how well large language models (LLMs) reproduce official crash coding. This study benchmarked six frontier LLMs by comparing narrative-derived crash attribute codes with corresponding fields in the Arkansas fatal-crash database. The analysis linked 5,587 fatal-crash narratives with 5,889 structured crash records from Arkansas (2015-2025), yielding 4,194 matched crashes. Six LLMs were evaluated using an identical zero-shot prompt to code crash manner, non-motorist relation, intersection type, work-zone relation, roadway surface condition, and light condition. Performance was evaluated using agreement, macro-averaged F1 score, Cohen's kappa, coverage, selective agreement, and comparisons with always-majority, always-Unknown, and keyword-rule baselines. Repeated-measures analyses and a generalized estimating equations model assessed differences among models and attributes. GPT-5.5 High achieved the highest agreement among the evaluated LLMs, but the always-majority baseline produced higher raw agreement and the keyword-rule baseline achieved macro-averaged F1 score and Cohen's kappa comparable to the best-performing LLM. Agreement was highest for non-motorist relation and crash manner and lowest for light condition, roadway surface condition, and work-zone relation. Differences across crash attributes exceeded differences across models. These results provide a benchmark for evaluating LLM-based crash coding and show that deployment should be evaluated on an attribute-specific basis using transparent baselines and human review.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。