构建复杂交通场景的视频问答数据集,提升智能交通系统中的推理能力。
InterAct-Video: Reasoning-Rich Video QA for Urban Traffic
- 设计跨时空维度的交通视频问答数据集,聚焦车辆交互与事件检测。
- 包含8小时真实路况、10秒片段、2.5万+问答对,覆盖多类交通动态。
- 适合研究智能交通、视频理解与多模态推理的学者和工程师使用。
交通监控对城市出行、道路安全和智能交通系统至关重要。深度学习通过视频问答(VideoQA)模型推动了基于视频的交通监控发展,实现了对交通视频的结构化信息提取。然而,现有视频问答模型在真实交通场景中面临挑战,因多重事件在时空维度上并发发生。为此,本文提出 extbf{InterAct VideoQA},一个专门用于评估和提升交通监控任务中视频问答模型性能的标注数据集。该数据集包含从多个路口采集的8小时真实交通视频,分割为10秒片段,涵盖超过25,000个问答对,内容涉及时空动态、车辆交互、事故检测及其他关键交通属性。对主流视频问答模型在该数据集上的评估揭示了其在复杂交通场景中细粒度时空依赖推理方面的不足。此外,模型在该数据集上微调后表现出显著性能提升,证明了领域专用数据集对视频问答的重要性。InterAct VideoQA 已公开,作为基准数据集支持未来可部署的智能交通视频问答模型研究。GitHub 仓库:https://github.com/joe-rabbit/InterAct_VideoQA
原文摘要 · Abstract (English)
Traffic monitoring is crucial for urban mobility, road safety, and intelligent transportation systems (ITS). Deep learning has advanced video-based traffic monitoring through video question answering (VideoQA) models, enabling structured insight extraction from traffic videos. However, existing VideoQA models struggle with the complexity of real-world traffic scenes, where multiple concurrent events unfold across spatiotemporal dimensions. To address these challenges, this paper introduces \textbf{InterAct VideoQA}, a curated dataset designed to benchmark and enhance VideoQA models for traffic monitoring tasks. The InterAct VideoQA dataset comprises 8 hours of real-world traffic footage collected from diverse intersections, segmented into 10-second video clips, with over 25,000 question-answer (QA) pairs covering spatiotemporal dynamics, vehicle interactions, incident detection, and other critical traffic attributes. State-of-the-art VideoQA models are evaluated on InterAct VideoQA, exposing challenges in reasoning over fine-grained spatiotemporal dependencies within complex traffic scenarios. Additionally, fine-tuning these models on InterAct VideoQA yields notable performance improvements, demonstrating the necessity of domain-specific datasets for VideoQA. InterAct VideoQA is publicly available as a benchmark dataset to facilitate future research in real-world deployable VideoQA models for intelligent transportation systems. GitHub Repo: https://github.com/joe-rabbit/InterAct_VideoQA
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。