评测顶尖视频问答模型在交通监控中的表现,发现其在复杂场景下仍有明显短板。
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks
- 用GPT-4o评估多个模型对交通视频的问答能力,涵盖检测、时序推理等任务。
- VideoLLaMA-2准确率达57%,在组合推理和答案一致性上表现最优。
- 当前模型在多目标跟踪与复杂场景理解上仍存不足,适合关注智能交通研究者参考。
视频问答(VideoQA)的最新进展为交通监控提供了广阔应用前景,其中高效视频理解至关重要。在智能交通系统(ITS)中,实时回答如“过去10分钟有多少辆红色汽车通过?”或“下午3:00至3:05之间是否发生事故?”等复杂问题,可显著提升态势感知与决策水平。尽管视觉语言模型取得进展,视频问答在动态环境中的多物体及复杂时空关系理解仍具挑战。本研究使用非基准的合成与真实交通视频序列,评估当前最先进的VideoQA模型。框架借助GPT-4o评估准确性、相关性与一致性,覆盖基础检测、时序推理与分解查询。VideoLLaMA-2表现最佳,准确率达57%,尤其在组合推理与答案一致性方面突出。然而,所有模型(包括VideoLLaMA-2)在多对象跟踪、时序连贯性与复杂场景解析方面仍存在局限,揭示了现有架构的不足。研究结果表明,VideoQA在交通监控中潜力巨大,但需提升多对象跟踪、时序推理与组合能力。改进这些方向将使其在事故检测、交通流管理与智能城市规划中发挥关键作用。研究代码与框架已开源:https://github.com/joe-rabbit/VideoQA_Pilot_Study
原文摘要 · Abstract (English)
Recent advances in video question answering (VideoQA) offer promising applications, especially in traffic monitoring, where efficient video interpretation is critical. Within ITS, answering complex, real-time queries like "How many red cars passed in the last 10 minutes?" or "Was there an incident between 3:00 PM and 3:05 PM?" enhances situational awareness and decision-making. Despite progress in vision-language models, VideoQA remains challenging, especially in dynamic environments involving multiple objects and intricate spatiotemporal relationships. This study evaluates state-of-the-art VideoQA models using non-benchmark synthetic and real-world traffic sequences. The framework leverages GPT-4o to assess accuracy, relevance, and consistency across basic detection, temporal reasoning, and decomposition queries. VideoLLaMA-2 excelled with 57% accuracy, particularly in compositional reasoning and consistent answers. However, all models, including VideoLLaMA-2, faced limitations in multi-object tracking, temporal coherence, and complex scene interpretation, highlighting gaps in current architectures. These findings underscore VideoQA's potential in traffic monitoring but also emphasize the need for improvements in multi-object tracking, temporal reasoning, and compositional capabilities. Enhancing these areas could make VideoQA indispensable for incident detection, traffic flow management, and responsive urban planning. The study's code and framework are open-sourced for further exploration: https://github.com/joe-rabbit/VideoQA_Pilot_Study
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。