arXiv:2602.21137cs.CV2026-02

构建城市交通视频问答数据集,推动多对象时空推理研究。

UDVideoQA: A Traffic Video Question Answering Dataset for Multi-Object Spatio-Temporal Reasoning in Urban Dynamics

  • 从16小时真实路况中提取数据,采用动态模糊保护隐私。
  • 含28000个问答对,平均每秒1个问题,覆盖多层推理任务。
  • 揭示视觉理解与因果推理之间的模型能力鸿沟,适合评估多模态系统。

理解城市交通中复杂的多智能体动态仍是视频语言模型面临的根本挑战。本文提出城市动态视频问答(UDVideoQA)基准数据集,捕捉真实城市场景中未预设的动态行为。数据集源自16小时不同城市交叉口在多样交通、天气和光照条件下的行车录像,采用事件驱动的动态模糊技术,在不牺牲场景保真度的前提下保障隐私。通过统一标注流程,生成8小时密集标注视频中的28,000个问答对,平均约每秒1个问题。其分类体系涵盖从基础理解到归因、事件推理、逆向推理及反事实推断的多层次推理层级,支持对视觉定位与因果推理能力的系统性评估。对10个SOTA视频语言模型在UDVideoQA上进行综合实验,并在互补的视频问题生成基准上测试8个模型。结果表明,模型普遍存在感知-推理差距:擅长抽象推理的模型常在基础视觉定位上表现不佳。尽管Gemini Pro零样本表现最佳,但将较小的Qwen2.5-VL 7B模型在UDVideoQA上微调后,性能可媲美专有系统。在VideoQGen任务中,Gemini 2.5 Pro、Qwen3 Max生成的问题最相关且复杂,但所有模型的语言多样性有限,凸显人工评估的重要性。UDVideoQA套件包括数据集、标注工具及面向VideoQA与VideoQGen的基准,为推进鲁棒、隐私友好的现实世界多模态推理提供基础。数据集地址:https://ud-videoqa.github.io/UD-VideoQA/UD-VideoQA/

原文摘要 · Abstract (English)

Understanding the complex, multi-agent dynamics of urban traffic remains a fundamental challenge for video language models. This paper introduces Urban Dynamics VideoQA, a benchmark dataset that captures the unscripted real-world behavior of dynamic urban scenes. UDVideoQA is curated from 16 hours of traffic footage recorded at multiple city intersections under diverse traffic, weather, and lighting conditions. It employs an event-driven dynamic blur technique to ensure privacy preservation without compromising scene fidelity. Using a unified annotation pipeline, the dataset contains 28K question-answer pairs generated across 8 hours of densely annotated video, averaging one question per second. Its taxonomy follows a hierarchical reasoning level, spanning basic understanding and attribution to event reasoning, reverse reasoning, and counterfactual inference, enabling systematic evaluation of both visual grounding and causal reasoning. Comprehensive experiments benchmark 10 SOTA VideoLMs on UDVideoQA and 8 models on a complementary video question generation benchmark. Results reveal a persistent perception-reasoning gap, showing models that excel in abstract inference often fail with fundamental visual grounding. While models like Gemini Pro achieve the highest zero-shot accuracy, fine-tuning the smaller Qwen2.5-VL 7B model on UDVideoQA bridges this gap, achieving performance comparable to proprietary systems. In VideoQGen, Gemini 2.5 Pro, and Qwen3 Max generate the most relevant and complex questions, though all models exhibit limited linguistic diversity, underscoring the need for human-centric evaluation. The UDVideoQA suite, including the dataset, annotation tools, and benchmarks for both VideoQA and VideoQGen, provides a foundation for advancing robust, privacy-aware, and real-world multimodal reasoning. UDVideoQA is available at https://ud-videoqa.github.io/UD-VideoQA/UD-VideoQA/.

视频问答城市交通因果推理多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。