arXiv:2412.07012cs.CVcs.AI2024-12被引 22

用程序化方法生成高质量多模态指令数据,提升模型理解图像能力。

ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models

  • 以场景图和人类编写程序为符号表示,程序化生成图像指令数据。
  • 构建超1000万条数据集,使模型在多个评测中性能提升最高达8%。
  • 适合做多模态模型训练的数据构建者与研究者使用。

随着多模态应用兴起,指令数据对训练能理解复杂图像查询的多模态语言模型至关重要。现有方法依赖强大但昂贵的大语言模型(LLMs)或多模态语言模型(MLMs)生成指令数据,易产生幻觉、存在版权问题,且难以扩展和解释。本文提出程序化方法 ProVision,利用场景图作为图像的符号表示,结合人工编写的程序系统合成视觉导向的指令数据。该方法确保生成过程可解释、可控,并高效扩展,同时保持事实准确性。通过实现24个单图像与14个多图像指令生成器及场景图生成流水线,构建可扩展、低成本系统ProVision,为任意图像生成涵盖物体、属性、关系、深度等的多样化问答对。应用于Visual Genome和DataComp数据集,生成超过1000万条指令数据(ProVision-10M),并用于多模态语言模型的预训练与指令微调阶段。在指令微调阶段采用单图像指令数据,在CVBench的2D与3D子集上分别取得最高7%和8%的提升,同时在QBench2、RealWorldQA、MMMU上提升3%;多图像指令数据在Mantis-Eval上带来8%改进。将数据同时用于xGen-MM-4B的预训练与微调,11个基准平均提升1.6%。

原文摘要 · Abstract (English)

With the rise of multimodal applications, instruction data has become critical for training multimodal language models capable of understanding complex image-based queries. Existing practices rely on powerful but costly large language models (LLMs) or multimodal language models (MLMs) to produce instruction data. These are often prone to hallucinations, licensing issues and the generation process is often hard to scale and interpret. In this work, we present a programmatic approach that employs scene graphs as symbolic representations of images and human-written programs to systematically synthesize vision-centric instruction data. Our approach ensures the interpretability and controllability of the data generation process and scales efficiently while maintaining factual accuracy. By implementing a suite of 24 single-image, 14 multi-image instruction generators, and a scene graph generation pipeline, we build a scalable, cost-effective system: ProVision which produces diverse question-answer pairs concerning objects, attributes, relations, depth, etc., for any given image. Applied to Visual Genome and DataComp datasets, we generate over 10 million instruction data points, ProVision-10M, and leverage them in both pretraining and instruction tuning stages of MLMs. When adopted in the instruction tuning stage, our single-image instruction data yields up to a 7% improvement on the 2D split and 8% on the 3D split of CVBench, along with a 3% increase in performance on QBench2, RealWorldQA, and MMMU. Our multi-image instruction data leads to an 8% improvement on Mantis-Eval. Incorporation of our data in both pre-training and fine-tuning stages of xGen-MM-4B leads to an averaged improvement of 1.6% across 11 benchmarks.

多模态指令数据程序生成视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。