构建交通场景时空理解新基准,统一问答、指代与定位任务。
TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes
- 用元组表达时空对象,统一三类任务的评估框架。
- 含1000段视频、8.5万道问答题,覆盖恶劣天气等复杂场景。
- 提出TUMTraffic-Qwen模型,揭示细粒度时空推理难点。
我们提出TUMTraffic-VideoQA,一个面向复杂路边交通场景的时空视频理解数据集与基准。数据集包含1,000段视频,涵盖85,000个多项选择问答对、2,300个物体描述和5,700个时空定位标注,覆盖恶劣天气与交通异常等多种真实条件。通过引入基于元组的时空对象表达,该基准统一了多项选择视频问答、参照物描述与时空定位三项核心任务。我们还提出了TUMTraffic-Qwen基线模型,采用视觉标记采样策略,为细粒度时空推理挑战提供洞见。大量实验验证了数据集的复杂性,揭示现有模型局限,确立其在智能交通系统研究中的坚实基础。数据集与基准已公开,供进一步探索。
原文摘要 · Abstract (English)
We present TUMTraffic-VideoQA, a novel dataset and benchmark designed for spatio-temporal video understanding in complex roadside traffic scenarios. The dataset comprises 1,000 videos, featuring 85,000 multiple-choice QA pairs, 2,300 object captioning, and 5,700 object grounding annotations, encompassing diverse real-world conditions such as adverse weather and traffic anomalies. By incorporating tuple-based spatio-temporal object expressions, TUMTraffic-VideoQA unifies three essential tasks-multiple-choice video question answering, referred object captioning, and spatio-temporal object grounding-within a cohesive evaluation framework. We further introduce the TUMTraffic-Qwen baseline model, enhanced with visual token sampling strategies, providing valuable insights into the challenges of fine-grained spatio-temporal reasoning. Extensive experiments demonstrate the dataset's complexity, highlight the limitations of existing models, and position TUMTraffic-VideoQA as a robust foundation for advancing research in intelligent transportation systems. The dataset and benchmark are publicly available to facilitate further exploration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。