构建48.9万条视觉推理数据,让多模态大模型像人一样一步步看图思考。
VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
- 构建包含48.9万例的多轮人类级视觉推理数据集,覆盖四大领域。
- 在VisReason-Pro上微调后,模型推理准确率显著提升,跨基准泛化能力增强。
- 适合研究多模态推理、智能体决策与可解释性系统的研究者使用。
链式思维(CoT)提示在大语言模型中已证明能有效激发复杂推理能力,但在多模态大语言模型(MLLMs)中的潜力仍待挖掘,主要受限于缺乏大规模、空间语义丰富的视觉推理数据集。现有视觉CoT资源通常规模小、领域专一,或缺乏人类般的逐步推理结构。本文提出VisReason,一个涵盖48.9万条标注样本的大规模数据集,覆盖四个不同领域,每条样本均包含多轮、类人的推理过程,引导MLLM进行可解释的视觉推理。在此基础上,我们进一步构建了16.5万条的VisReason-Pro子集,采用更高级别的专家级GPT标注器生成,包含详细的推理轨迹和基于深度信息的3D空间定位。在状态前沿的Qwen2.5-VL模型上微调,结果显示其在步骤式视觉推理准确性、可解释性及跨基准泛化能力上均有显著提升。这些结果表明,VisReason使MLLM具备更系统、更具泛化性的推理能力。我们期待其成为培养类人视觉推理的核心基础,推动下一代多模态智能的发展。
原文摘要 · Abstract (English)
Chain-of-Thought (CoT) prompting has proven remarkably effective for eliciting complex reasoning in large language models (LLMs). Yet, its potential in multimodal large language models (MLLMs) remains largely untapped, hindered by the absence of large-scale datasets that capture the rich, spatially grounded reasoning intrinsic to visual understanding. Existing visual-CoT resources are typically small, domain-specific, or lack the human-like stepwise structure necessary for compositional visual reasoning. In this paper, we introduce VisReason, a large-scale dataset designed to advance visual Chain-of-Thought reasoning. VisReason comprises 489K annotated examples spanning four diverse domains, each featuring multi-round, human-like rationales that guide MLLMs through interpretable visual reasoning steps. Building upon this, we curate VisReason-Pro, a 165K subset produced with a stronger expert-level GPT annotator, enriched with detailed reasoning traces and 3D spatial grounding via depth-informed annotations. Fine-tuning the state-of-the-art Qwen2.5-VL model on VisReason and VisReason-Pro yields substantial improvements in step-by-step visual reasoning accuracy, interpretability, and cross-benchmark generalization. These results demonstrate that VisReason equips MLLMs with more systematic and generalizable reasoning capabilities. We envision VisReason as a cornerstone for cultivating human-like visual reasoning, paving the way toward the next generation of multimodal intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。