arXiv:2410.02740cs.CVcs.AI2024-10被引 15

混合使用合成与原始图文描述能更好提升多模态模型性能

Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models

  • 设计可控制的流水线生成适配不同模型的多种描述格式
  • 保留原始AltText与合成描述的混合策略表现更优
  • 揭示各类模型对描述格式的独特偏好,指导训练优化

近期多模态模型的发展表明重写描述能提升性能,但关键问题仍存。尽管合成描述通常质量更高、图文对齐更好,但其是否可完全替代原始网页爬取的AltText尚不明确。不同多模态基础模型对描述格式可能存在独特偏好,但尚未有系统研究。本文提出一种新型可控、可扩展的描述生成流水线,针对CLIP、多模态大语言模型及扩散模型等,系统研究短合成描述(SSC)到密集合成描述(DSC+)的影响及其与AltText的交互。结果表明,保留合成描述与AltText的混合策略优于仅用合成描述,显著提升对齐度与模型性能,且各模型对特定描述格式有明显偏好。本研究为优化多模态预训练中的描述策略提供了重要依据。

原文摘要 · Abstract (English)

Recent advancements in multimodal models highlight the value of rewritten captions for improving performance, yet key challenges remain. For example, while synthetic captions often provide superior quality and image-text alignment, it is not clear whether they can fully replace AltTexts: the role of synthetic captions and their interaction with original web-crawled AltTexts in pre-training is still not well understood. Moreover, different multimodal foundation models may have unique preferences for specific caption formats, but efforts to identify the optimal captions for each model remain limited. In this work, we propose a novel, controllable, and scalable captioning pipeline designed to generate diverse caption formats tailored to various multimodal models. By examining Short Synthetic Captions (SSC) towards Dense Synthetic Captions (DSC+) as case studies, we systematically explore their effects and interactions with AltTexts across models such as CLIP, multimodal LLMs, and diffusion models. Our findings reveal that a hybrid approach that keeps both synthetic captions and AltTexts can outperform the use of synthetic captions alone, improving both alignment and performance, with each model demonstrating preferences for particular caption formats. This comprehensive analysis provides valuable insights into optimizing captioning strategies, thereby advancing the pre-training of multimodal foundation models.

多模态图像描述预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。