用1200万条带推理链的多模态指令数据,提升开源模型的复杂推理能力。
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
- 构建1200万条带中间推理过程的多模态指令数据,支持复杂任务训练。
- 在MathVerse、MMMU-Pro等基准上分别提升8.1%、7%和13.3%。
- 适合需要强多模态推理能力的研究者与开发者使用。
开源多模态大语言模型在多种任务中展现出巨大潜力,但其推理能力受限于现有指令微调数据集,这些数据集主要源自VQA、AI2D、ChartQA等学术数据集,任务简单且仅提供短语级答案,缺乏中间推理过程。为此,我们提出一种可扩展、低成本的方法,构建大规模多模态指令微调数据集,包含丰富中间推理链,旨在激发思维链(CoT)推理能力。仅使用开源模型,我们生成了包含1200万条指令-响应对的数据集,覆盖多样化的高推理强度任务,并确保推理过程详尽且真实。实验表明,基于该数据集训练的多模态大模型显著提升推理能力,在MathVerse(+8.1%)、MMMU-Pro(+7%)和MuirBench(+13.3%)等基准上达到顶尖水平。此外,非推理类基准也获得最高4%的性能提升。消融研究进一步验证了重写和自过滤等关键组件在数据构建中的重要性。
原文摘要 · Abstract (English)
Open-source multimodal large language models (MLLMs) have shown significant potential in a broad range of multimodal tasks. However, their reasoning capabilities remain constrained by existing instruction-tuning datasets, which were predominately repurposed from academic datasets such as VQA, AI2D, and ChartQA. These datasets target simplistic tasks, and only provide phrase-level answers without any intermediate rationales. To address these challenges, we introduce a scalable and cost-effective method to construct a large-scale multimodal instruction-tuning dataset with rich intermediate rationales designed to elicit CoT reasoning. Using only open models, we create a dataset containing 12M instruction-response pairs to cover diverse, reasoning-intensive tasks with detailed and faithful rationales. Experiments demonstrate that training MLLMs on this dataset significantly improves reasoning capabilities, achieving state-of-the-art performance on benchmarks such as MathVerse (+8.1%), MMMU-Pro (+7%), and MuirBench (+13.3%). Additionally, the model demonstrates notable improvements of up to 4% on non-reasoning-based benchmarks. Ablation studies further highlight the importance of key components, such as rewriting and self-filtering, in the dataset construction process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。