arXiv:2510.01582cs.CVcs.LG2025-10被引 1

构建25万张带思维链的图文数据集,助力视觉语言模型推理能力提升。

ImageNet-Think-250K: A Large-Scale Synthetic Dataset for Multimodal Reasoning for Vision Language Models

  • 用两款先进视觉语言模型生成带思维链的图文对。
  • 每张图配两组思考-答案序列,共25万条数据。
  • 适合研究多模态推理机制与模型训练的学者使用。

我们构建了ImageNet-Think,一个旨在提升视觉语言模型(VLMs)显式推理能力的多模态推理数据集。该数据集基于ImageNet21k中的25万张图像,提供结构化的思维链标记和对应答案。数据由两款先进的VLMs——GLM-4.1V-9B-Thinking和Kimi-VL-A3B-Thinking-2506生成。每张图像配有两组思考-答案序列,用于训练与评估多模态推理模型。本数据集捕捉了VLMs的逐步推理过程及最终描述性答案。我们的目标是推动更鲁棒的VLMs发展,并增进对多模态推理机制的理解。数据集与评估基准将公开发布,以支持相关研究。

原文摘要 · Abstract (English)

We develop ImageNet-Think, a multimodal reasoning dataset designed to aid the development of Vision Language Models (VLMs) with explicit reasoning capabilities. Our dataset is built on 250,000 images from ImageNet21k dataset, providing structured thinking tokens and corresponding answers. Our synthetic dataset is generated by two state-of-the-art VLMs: GLM-4.1V-9B-Thinking and Kimi-VL-A3B-Thinking-2506. Each image is accompanied by two pairs of thinking-answer sequences, creating a resource for training and evaluating multimodal reasoning models. We capture the step-by-step reasoning process of VLMs and the final descriptive answers. Our goal with this dataset is to enable the development of more robust VLMs while contributing to the broader understanding of multimodal reasoning mechanisms. The dataset and evaluation benchmarks will be publicly available to aid research in reasoning/thinking multimodal VLMs.

多模态推理思维链视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。