用合成图文对提升大模型对复合图像的理解能力
CompCap: Improving Multimodal Large Language Models with Composite Captions
- 用大模型自动构建11.8万组复合图像与精准描述的配对数据
- 在11个评测基准上使模型性能平均提升1.7%至2.9%
- 特别适合提升模型处理图表、海报等合成视觉内容的能力
多模态大模型(MLLMs)对复合图像(CIs)的理解能力如何?复合图像由图表、海报或截图等元素拼接而成,而非真实拍摄。尽管在实际应用中广泛存在,现有研究仍主要聚焦于自然图像(NIs)。我们发现当前MLLMs在理解CIs时面临显著挑战,难以准确提取信息或进行复杂推理。现有CI训练数据多用于问答任务(如ChartQA、ScienceQA),而高质量的图像-文本对数据集仅存在于NIs。为此,我们提出复合图像描述框架CompCap,利用大语言模型与自动化工具生成带精确描述的复合图像。基于此,我们构建了包含11.8万组图像-文本对的CompCap-118K数据集,涵盖六类复合图像。通过监督微调三种规模的MLLMs(xGen-MM-inst.-4B、LLaVA-NeXT-Vicuna-7B/13B),实验表明,CompCap-118K在11个基准上分别带来1.7%、2.0%和2.9%的平均性能提升。
原文摘要 · Abstract (English)
How well can Multimodal Large Language Models (MLLMs) understand composite images? Composite images (CIs) are synthetic visuals created by merging multiple visual elements, such as charts, posters, or screenshots, rather than being captured directly by a camera. While CIs are prevalent in real-world applications, recent MLLM developments have primarily focused on interpreting natural images (NIs). Our research reveals that current MLLMs face significant challenges in accurately understanding CIs, often struggling to extract information or perform complex reasoning based on these images. We find that existing training data for CIs are mostly formatted for question-answer tasks (e.g., in datasets like ChartQA and ScienceQA), while high-quality image-caption datasets, critical for robust vision-language alignment, are only available for NIs. To bridge this gap, we introduce Composite Captions (CompCap), a flexible framework that leverages Large Language Models (LLMs) and automation tools to synthesize CIs with accurate and detailed captions. Using CompCap, we curate CompCap-118K, a dataset containing 118K image-caption pairs across six CI types. We validate the effectiveness of CompCap-118K by supervised fine-tuning MLLMs of three sizes: xGen-MM-inst.-4B and LLaVA-NeXT-Vicuna-7B/13B. Empirical results show that CompCap-118K significantly enhances MLLMs' understanding of CIs, yielding average gains of 1.7%, 2.0%, and 2.9% across eleven benchmarks, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。