用AI自动生成交通事故图,提升交通分析效率
Automating Crash Diagram Generation Using Vision-Language Models: A Case Study on Multi-Lane Roundabouts

- 设计三段式提示框架引导模型理解、提取、绘图
- GPT-4o生成图平均得分6.29/10,优于其他模型
- 适合交通工程、城市规划人员快速辅助分析事故
事故图是交通安全管理的重要工具,但人工绘制耗时且易出错。本研究探索利用视觉语言模型(VLMs)从警察事故报告中自动化生成多车道环形交叉口的事故图,作为高难度测试案例。提出三阶段结构化提示框架,引导模型完成解读、信息提取与视觉合成,并设计包含10项指标的评估体系,衡量语义准确性、空间保真度与视觉清晰度。在79份事故报告上测试了GPT-4o、Gemini-1.5-Flash和Janus-4o三个主流模型,GPT-4o表现最佳,平均得分6.29(满分10),次为Gemini-1.5-Flash(5.28)和Janus-4o(3.64)。分析显示GPT-4o在空间推理与数据一致性方面优势明显。结果表明VLMs在工程可视化任务中潜力巨大,但仍存局限。该研究为将生成式AI融入事故分析流程奠定基础,有望提升效率、一致性和可解释性。
原文摘要 · Abstract (English)
Crash diagrams are essential tools in transportation safety analysis, yet their manual preparation remains time-consuming and prone to human variability. This study investigates the use of Vision-Language Models (VLMs) to automate crash diagram generation from police crash reports, focusing on multilane roundabouts as a challenging test case. A three-part structured prompt framework was developed to guide model reasoning through interpretation, extraction, and visual synthesis, while a 10-metric evaluation system was designed to assess diagram quality in terms of semantic accuracy, spatial fidelity, and visual clarity. Three popular models, including GPT-4o, Gemini-1.5-Flash, and Janus-4o, were tested on 79 crash reports. GPT-4o achieved the highest average performance (6.29 out of 10), followed by Gemini-1.5-Flash (5.28) and Janus-4o (3.64). The analysis revealed GPT-4o's superior spatial reasoning and alignment between extracted and visualized crash data. These results highlight both the promise and current limitations of VLMs in engineering visualization tasks. The study lays the groundwork for integrating generative AI into crash analysis workflows to improve efficiency, consistency, and interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。