早加图像数据能提升视觉语言模型表现,且不影响文本任务能力。
Should VLMs be Pre-trained with Image Data?
- 在预训练中混合图文数据,比后期加图更有效
- 10亿参数模型在80%预训练阶段引入视觉信息,平均性能提升2%
- 兼顾图文与纯文本任务,适合多模态应用
仅用文本预训练的模型若在后续阶段加入图像数据,可在视觉语言任务上表现良好。然而,这种两阶段方法相比从早期就融合图像的视觉语言模型(VLM)究竟带来多少增益或损失尚不明确。为此,我们训练了涵盖多种数据集、规模、图文比例及不同预训练阶段引入视觉标记的模型,并在一系列视觉语言和纯文本任务上进行微调与评估。结果表明,在预训练中混合图文数据,可使模型在视觉语言任务上表现更优,同时保持对纯文本任务的强适应性。在6个多样化任务上的平均表现显示,对于10亿参数模型,若在预训练完成80%时引入视觉标记,相比完全预训练后再引入,平均性能提升2%。
原文摘要 · Abstract (English)
Pre-trained LLMs that are further trained with image data perform well on vision-language tasks. While adding images during a second training phase effectively unlocks this capability, it is unclear how much of a gain or loss this two-step pipeline gives over VLMs which integrate images earlier into the training process. To investigate this, we train models spanning various datasets, scales, image-text ratios, and amount of pre-training done before introducing vision tokens. We then fine-tune these models and evaluate their downstream performance on a suite of vision-language and text-only tasks. We find that pre-training with a mixture of image and text data allows models to perform better on vision-language tasks while maintaining strong performance on text-only evaluations. On an average of 6 diverse tasks, we find that for a 1B model, introducing visual tokens 80% of the way through pre-training results in a 2% average improvement over introducing visual tokens to a fully pre-trained model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。