arXiv:2509.09730cs.CVcs.AI2025-09中稿 · Image and Vision C…被引 4

首个面向智能交通监控的多模态大模型数据集,显著提升视觉语言模型表现。

MITS: A Large-Scale Multimodal Benchmark Dataset for Intelligent Traffic Surveillance

  • 构建17万张真实交通监控图像,涵盖24个子类目标与事件
  • 生成500万条指令跟随问答对,覆盖五大核心交通任务
  • 开源数据集与模型,助力智能交通系统研发

通用多模态大模型在图像文本任务中取得显著进展,但在智能交通监控(ITS)领域表现仍受限,主要因缺乏专用多模态数据集。为此,我们提出MITS(Multimodal Intelligent Traffic Surveillance),首个专为ITS设计的大规模多模态基准数据集。MITS包含170,400张来自交通监控摄像头的真实世界图像,涵盖8大类、24个子类的ITS特定目标与事件,在多样环境条件下进行标注。通过系统化数据生成流程,构建高质量图像描述及500万条指令跟随式视觉问答对,覆盖五大关键ITS任务:目标与事件识别、目标计数、目标定位、背景分析与事件推理。为验证其有效性,我们在该数据集上微调主流多模态大模型,推动开发专用应用。实验表明,经MITS微调后,LLaVA-1.5性能从0.494升至0.905(+83.2%),LLaVA-1.6从0.678增至0.921(+35.8%),Qwen2-VL从0.584升至0.926(+58.6%),Qwen2.5-VL从0.732升至0.930(+27.0%)。数据集、代码与模型均已开源,为推进ITS与多模态大模型研究提供高价值资源。

原文摘要 · Abstract (English)

General-domain large multimodal models (LMMs) have achieved significant advances in various image-text tasks. However, their performance in the Intelligent Traffic Surveillance (ITS) domain remains limited due to the absence of dedicated multimodal datasets. To address this gap, we introduce MITS (Multimodal Intelligent Traffic Surveillance), the first large-scale multimodal benchmark dataset specifically designed for ITS. MITS includes 170,400 independently collected real-world ITS images sourced from traffic surveillance cameras, annotated with eight main categories and 24 subcategories of ITS-specific objects and events under diverse environmental conditions. Additionally, through a systematic data generation pipeline, we generate high-quality image captions and 5 million instruction-following visual question-answer pairs, addressing five critical ITS tasks: object and event recognition, object counting, object localization, background analysis, and event reasoning. To demonstrate MITS's effectiveness, we fine-tune mainstream LMMs on this dataset, enabling the development of ITS-specific applications. Experimental results show that MITS significantly improves LMM performance in ITS applications, increasing LLaVA-1.5's performance from 0.494 to 0.905 (+83.2%), LLaVA-1.6's from 0.678 to 0.921 (+35.8%), Qwen2-VL's from 0.584 to 0.926 (+58.6%), and Qwen2.5-VL's from 0.732 to 0.930 (+27.0%). We release the dataset, code, and models as open-source, providing high-value resources to advance both ITS and LMM research.

智能交通多模态数据集视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。