arXiv:2411.07461cs.CVcs.AI2024-11被引 12

构建2.18亿图文对数据集,让图像描述更准确真实。

BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions

  • 用大规模网络文本增强合成图像描述,提升事实准确性。
  • 在21800万图文对上训练模型,显著提升多模态任务表现。
  • 适合想训练更懂知识的视觉语言模型的研究者使用。

我们提出BLIP3-KALE,一个包含2.18亿张图像与文本配对的数据集,弥合了描述性合成标题与真实网络级替代文本之间的差距。KALE通过引入网络规模的替代文本,对合成密集图像描述进行知识增强,生成具有事实依据的图像描述。采用两阶段方法,利用大型视觉语言模型和语言模型生成知识增强型描述,并以此训练专用视觉语言模型以扩展数据集规模。在KALE上训练的视觉语言模型在多项多模态任务中表现更优,实验验证了其在训练更具能力与知识的多模态模型方面的有效性。KALE数据集已发布于https://huggingface.co/datasets/Salesforce/blip3-kale。

原文摘要 · Abstract (English)

We introduce BLIP3-KALE, a dataset of 218 million image-text pairs that bridges the gap between descriptive synthetic captions and factual web-scale alt-text. KALE augments synthetic dense image captions with web-scale alt-text to generate factually grounded image captions. Our two-stage approach leverages large vision-language models and language models to create knowledge-augmented captions, which are then used to train a specialized VLM for scaling up the dataset. We train vision-language models on KALE and demonstrate improvements on vision-language tasks. Our experiments show the utility of KALE for training more capable and knowledgeable multimodal models. We release the KALE dataset at https://huggingface.co/datasets/Salesforce/blip3-kale

多模态图像描述数据集知识增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。