用低幻觉合成图文对训练视觉语言模型,效果优于真实数据。
Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training
- 通过连续DPO方法生成高质量、低幻觉的合成图文对。
- 合成数据使模型在15项任务上提升6.2%以上,认知任务超alt-text 7.5%。
- 适合大规模预训练场景,尤其对数据稀缺领域有显著帮助。
近年来,视觉语言模型预训练取得快速发展,主要得益于大语言模型文本能力的持续提升。然而,现有多模态大模型训练严重依赖高质量图像-文本配对数据。随着模型与数据规模指数级增长,精心标注的数据日益稀缺且饱和,严重制约该领域进一步发展。本研究探索了适用于视觉语言模型预训练的大规模合成图文生成技术,实证表明:大规模低幻觉合成图文对可双重作用——一是作为真实数据的有效替代,二是集成后能显著提升模型性能。本文提出一种新颖的合成图文生成流水线,结合连续DPO方法,在70亿参数模型上,测试集无幻觉图文率从48.3%提升至77.9%。实证验证显示,使用本方法生成数据训练的模型,在15个视觉语言任务中相比使用alt-text的相同图像,性能提升至少6.2%;在20个常见认知领域,表现优于alt-text至少7.5%。此外,在文生图任务中也表现优异:在真实世界验证集上FID降低17.1,在MSCOCO验证集上降低13.3。
原文摘要 · Abstract (English)
In recent years, the field of vision-language model pre-training has experienced rapid advancements, driven primarily by the continuous enhancement of textual capabilities in large language models. However, existing training paradigms for multimodal large language models heavily rely on high-quality image-text pairs. As models and data scales grow exponentially, the availability of such meticulously curated data has become increasingly scarce and saturated, thereby severely limiting further advancements in this domain. This study investigates scalable caption generation techniques for vision-language model pre-training and demonstrates that large-scale low-hallucination synthetic captions can serve dual purposes: 1) acting as a viable alternative to real-world data for pre-training paradigms and 2) achieving superior performance enhancement when integrated into vision-language models through empirical validation. This paper presents following key contributions: 1) a novel pipeline for generating high-quality, low-hallucination, and knowledge-rich synthetic captions. Our continuous DPO methodology yields remarkable results in reducing hallucinations. Specifically, the non-hallucination caption rate on a held-out test set increases from 48.3% to 77.9% for a 7B-size model. 2) Comprehensive empirical validation reveals that our synthetic captions confer superior pre-training advantages over their counterparts. Across 15 vision language tasks, the model trained with our data achieves a significant performance gain of at least 6.2% compared to identical images with alt-text. In 20 common cognitive domains, the model trained with our data outperforms the alt-text data by at least 7.5%. Meanwhile, it also offers considerable support in the text-to-image domain. With our dataset, the FID score is reduced by 17.1 on a real-world validation benchmark and 13.3 on the MSCOCO validation benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。