用视觉语言模型生成2600万对多模态数据,大幅提升检索性能。
MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval
- 利用开源视觉语言模型和开放图像合成海量多模态数据
- 在4个基准上达成当前最佳零样本性能,36个MMEB数据集表现最优
- 适合追求高效训练与高泛化能力的多模态研究者
尽管多模态检索需求快速增长,但进展仍受限于训练数据不足。本文提出MegaPairs,一种基于视觉语言模型(VLMs)和开放域图像的数据合成方法,并构建了大规模合成数据集。实证分析表明,MegaPairs生成的数据质量高,使多模态检索器性能显著超越在70倍更多现有数据上训练的基线模型。由于仅依赖通用图像语料和开源VLMs,该方法可轻松扩展,实现检索性能持续提升。本阶段已生成超过2600万条训练实例,并训练了多个不同规模的模型。这些模型在4个主流组合图像检索(CIR)基准上达到当前最优零样本性能,在MMEB提供的36个数据集上综合表现最高,且通过下游微调进一步提升效果。所产数据集、预训练模型及合成流水线将公开,推动该领域发展。
原文摘要 · Abstract (English)
Despite the rapidly growing demand for multimodal retrieval, progress in this field remains severely constrained by a lack of training data. In this paper, we introduce MegaPairs, a novel data synthesis method that leverages vision language models (VLMs) and open-domain images, together with a massive synthetic dataset generated from this method. Our empirical analysis shows that MegaPairs generates high-quality data, enabling the multimodal retriever to significantly outperform the baseline model trained on 70$\times$ more data from existing datasets. Moreover, since MegaPairs solely relies on general image corpora and open-source VLMs, it can be easily scaled up, enabling continuous improvements in retrieval performance. In this stage, we produced more than 26 million training instances and trained several models of varying sizes using this data. These new models achieve state-of-the-art zero-shot performance across 4 popular composed image retrieval (CIR) benchmarks and the highest overall performance on the 36 datasets provided by MMEB. They also demonstrate notable performance improvements with additional downstream fine-tuning. Our produced dataset, well-trained models, and data synthesis pipeline will be made publicly available to facilitate the future development of this field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。