构建4000万条高质量多模态指令数据,提升视觉语言模型性能。
Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data
- 通过统一预处理与合成生成,构建超大规模多模态指令数据集。
- 基于该数据训练的20亿参数模型在同类中达最新水平。
- 适合研究多模态大模型、数据增强与开源模型优化的学者使用。
近期,视觉语言模型(VLMs)在多模态任务中取得显著进展,而多模态指令数据是提升其能力的基础。尽管已有多个开源多模态数据集,但公开指令数据在规模和质量上的局限性,导致基于这些数据训练的VLM性能远落后于闭源模型。为此,我们提出Infinity-MM,一个大规模多模态指令数据集。我们整合现有数据并进行统一预处理,获得超过4000万条样本,保证多样性与准确性。为进一步支持数据的规模化扩展与高质量持续获取,我们提出一种基于标签系统的合成指令生成方法,利用开源VLM和图像-指令类型对应关系实现高效数据生成。基于此高质量数据,我们训练了20亿参数的视觉语言模型Aquila-VL-2B,其在同规模模型中达到最先进性能。数据已开放:https://huggingface.co/datasets/BAAI/Infinity-MM。
原文摘要 · Abstract (English)
Recently, Vision-Language Models (VLMs) have achieved remarkable progress in multimodal tasks, and multimodal instruction data serves as the foundation for enhancing VLM capabilities. Despite the availability of several open-source multimodal datasets, limitations in the scale and quality of open-source instruction data hinder the performance of VLMs trained on these datasets, leading to a significant gap compared to models trained on closed-source data. To address this challenge, we introduce Infinity-MM, a large-scale multimodal instruction dataset. We collected the available multimodal instruction datasets and performed unified preprocessing, resulting in a dataset with over 40 million samples that ensures diversity and accuracy. Furthermore, to enable large-scale expansion of instruction data and support the continuous acquisition of high-quality data, we propose a synthetic instruction generation method based on a tagging system and open-source VLMs. By establishing correspondences between different types of images and associated instruction types, this method can provide essential guidance during data synthesis. Leveraging this high-quality data, we have trained a 2-billion-parameter Vision-Language Model, Aquila-VL-2B, which achieves state-of-the-art (SOTA) performance among models of similar scale. The data is available at: https://huggingface.co/datasets/BAAI/Infinity-MM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。