arXiv:2606.21419cs.CVcs.AI2026-06

构建大规模跨域图文数据集,支持细粒度视觉语言学习

MIRCaps: A Large-Scale Mixed-Domain Dataset with Image-Level and Region-Level Captions for Fine-Grained Vision-Language Learning

论文配图:MIRCaps: A Large-Scale Mixed-Domain Dataset with Image-Level and Region-Level Captions for Fine-Grained Vision-Language Learning
图 1 · 摘自论文原文
  • 包含图像级与区域级双类型标注,每图平均7个描述
  • 覆盖14万张图像,170万条区域描述,提升模型感知细节能力
  • 适合做视觉定位、细粒度图像生成的研究者使用

尽管视觉语言模型取得进展,但面向通用场景与监控视频系统的跨域图像-文本数据集仍十分有限。为此,我们构建了一个大规模多模态数据集,包含141,364张图像、981,947条图像级描述、1,742,264条区域级描述及1,391,779个边界框标注。每张图像平均配有七条描述整体场景的图像级标题,以及每个标注框对应七条区域级描述。这些互补的标注形式旨在帮助模型学习对象类别、估计尺寸、颜色、动作、状态及环境上下文等细粒度视觉属性。我们在图像描述生成和目标检测两个下游任务上验证了该数据集的有效性。实验表明,SmolVLM-256M-Instruct、BLIP、BLIP2和Qwen2.5-VL 3B-Instruct等轻量级视觉语言模型均可通过本数据集有效微调。数据集与代码已公开于https://zenodo.org/records/20418601。

原文摘要 · Abstract (English)

Despite recent progress in Vision-Language Models (VLMs), mixed-domain image-caption datasets for both general-purpose and CCTV-based video surveillance systems remain limited. To address this gap, we introduce a large-scale multimodal dataset comprising 141,364 images, 981,947 image-level captions, 1,742,264 region-level captions, and 1,391,779 bounding box annotations. Each image is associated with an average of seven image-level captions describing different aspects of the overall scene, as well as seven region-level captions for each annotated bounding box. These complementary caption types are designed to help VLMs learn fine-grained visual attributes, including object categories, estimated sizes, colors, actions, states, and surrounding environmental context. We demonstrate the effectiveness of the dataset on two important downstream tasks: image captioning and object detection. Experimental results show that lightweight VLMs, including SmolVLM-256M-Instruct, BLIP, BLIP2, and Qwen2.5-VL 3B-Instruct, can be effectively fine-tuned using our dataset. Our dataset and code are publicly available at https://zenodo.org/records/20418601.

多模态数据集细粒度理解视觉语言模型图像标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。