用GPT-4o分析鸟瞰视频,自动识别路口冲突并给出解释。
Visual Reasoning at Urban Intersections: FineTuning GPT-4o for Traffic Conflict Detection
- 用GPT-4o直接处理鸟瞰视角视频,结合视觉与逻辑推理
- 模型检测准确率达77.14%,解释与建议准确率分别达89.9%和92.3%
- 适合智能交通系统开发、城市路口安全优化研究者参考
无信号控制的城市场口因结构复杂、冲突频繁且存在盲区,管理难度大。本研究探索利用多模态大语言模型(MLLMs),如GPT-4o,通过直接输入四路路口的鸟瞰视频实现视觉与逻辑推理。所提方法中,GPT-4o作为智能系统,可检测交通冲突,并生成解释与驾驶建议。微调后的模型检测准确率达77.14%;人工评估显示,模型生成的解释准确率为89.9%,推荐下一步动作准确率为92.3%。结果表明,基于视频输入的MLLM在实时交通管理中具有可行性,能为路口管理与运营提供可扩展、可操作的洞察。代码已开源:https://github.com/sarimasri3/Traffic-Intersection-Conflict-Detection-using-images.git。
原文摘要 · Abstract (English)
Traffic control in unsignalized urban intersections presents significant challenges due to the complexity, frequent conflicts, and blind spots. This study explores the capability of leveraging Multimodal Large Language Models (MLLMs), such as GPT-4o, to provide logical and visual reasoning by directly using birds-eye-view videos of four-legged intersections. In this proposed method, GPT-4o acts as intelligent system to detect conflicts and provide explanations and recommendations for the drivers. The fine-tuned model achieved an accuracy of 77.14%, while the manual evaluation of the true predicted values of the fine-tuned GPT-4o showed significant achievements of 89.9% accuracy for model-generated explanations and 92.3% for the recommended next actions. These results highlight the feasibility of using MLLMs for real-time traffic management using videos as inputs, offering scalable and actionable insights into intersections traffic management and operation. Code used in this study is available at https://github.com/sarimasri3/Traffic-Intersection-Conflict-Detection-using-images.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。