arXiv:2503.21459cs.CV2025-03CVPR被引 19

构建首个覆盖全球道路事件的社交视频问答数据集,助力智能交通理解

RoadSocial: A Diverse VideoQA Dataset and Benchmark for Road Event Understanding from Social Video Narratives

  • 基于社交媒体视频与评论,用大模型半自动生成多样问答对
  • 涵盖13.2万视频、26万高质量问答对,覆盖12类复杂任务
  • 适合研究通用视频大模型在真实道路场景下的理解能力

我们提出RoadSocial,一个大规模、多样化的视频问答数据集,专用于从社交媒体叙述中理解通用道路事件。与受限于区域偏差、视角偏差和专家标注的现有数据集不同,RoadSocial通过多地理区域、多摄像机视角(监控、手持、无人机)和丰富的社交讨论,捕捉道路事件的全球复杂性。我们的可扩展半自动标注框架利用文本大模型和视频大模型,在12个具有挑战性的问答任务上生成全面的问答对,推动道路事件理解的边界。RoadSocial源自覆盖1400万帧图像和41.4万条社交评论的社交媒体视频,最终形成包含13.2万视频、674个标签和26万条高质量问答对的数据集。我们在该基准上评估了18种视频大模型(开源与专有、驾驶专用与通用),并证明其能有效提升通用视频大模型的道路事件理解能力。

原文摘要 · Abstract (English)

We introduce RoadSocial, a large-scale, diverse VideoQA dataset tailored for generic road event understanding from social media narratives. Unlike existing datasets limited by regional bias, viewpoint bias and expert-driven annotations, RoadSocial captures the global complexity of road events with varied geographies, camera viewpoints (CCTV, handheld, drones) and rich social discourse. Our scalable semi-automatic annotation framework leverages Text LLMs and Video LLMs to generate comprehensive question-answer pairs across 12 challenging QA tasks, pushing the boundaries of road event understanding. RoadSocial is derived from social media videos spanning 14M frames and 414K social comments, resulting in a dataset with 13.2K videos, 674 tags and 260K high-quality QA pairs. We evaluate 18 Video LLMs (open-source and proprietary, driving-specific and general-purpose) on our road event understanding benchmark. We also demonstrate RoadSocial's utility in improving road event understanding capabilities of general-purpose Video LLMs.

视频问答道路理解社交视频数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。