用细粒度视觉标注提升多模态大模型理解与生成能力
Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models
- 构建包含粗细粒度概念标注的多模态数据集MMGiC
- 结合细粒度标注后在多个任务上提升3.95%和2.34%准确率
- 适合研究多模态对齐与视觉概念理解的学者使用
多模态大语言模型(MLLMs)在视觉-语言任务中表现优异,主要依赖粗粒度概念标注(如图像描述)进行预训练。我们提出假设:引入细粒度概念标注(如物体标签与区域)可进一步提升性能,因不同粒度的标注在概念表征的广度与深度上具有互补性。为此,我们构建了名为MMGiC的新数据集,包含多粒度多模态概念标注。通过探索不同数据配方对多模态理解与生成的影响,分析表明,在结构化模板和通用MLLM框架下,多粒度标注能有效融合与互补。实验验证了其帮助模型更精准定位与学习概念的能力,实现视觉与语言在多粒度上的对齐。在12个基准任务上,结合MMGiC与图像-描述数据相比仅用后者时,分别在POPE和SEED-Bench上取得3.95%和2.34%的绝对性能提升。代码、数据与模型将开源于https://github.com/LooperXX/MMGiC。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) excel in vision--language tasks by pre-training solely on coarse-grained concept annotations (e.g., image captions). We hypothesize that integrating fine-grained concept annotations (e.g., object labels and object regions) will further improve performance, as both data granularities complement each other in terms of breadth and depth in concept representation. We introduce a new dataset featuring Multimodal Multi-Grained Concept annotations (MMGiC) for MLLMs. In constructing MMGiC, we explore the impact of different data recipes on multimodal comprehension and generation. Our analyses reveal that multi-grained concept annotations integrate and complement each other, under our structured template and a general MLLM framework. We clearly explore and demonstrate the potential of MMGiC to help MLLMs better locate and learn concepts, aligning vision and language at multiple granularities. We further validate our hypothesis by investigating the fair comparison and effective collaboration between MMGiC and image--caption data on 12 multimodal comprehension and generation benchmarks, e.g., their appropriate combination achieve 3.95% and 2.34% absolute improvements over image--caption data alone on POPE and SEED-Bench. Code, data and models will be available at https://github.com/LooperXX/MMGiC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。