用多专家协作生成高质量图像标注,提升视觉语言模型理解力
FullAnno: A Data Engine for Enhancing Image Comprehension of MLLMs
- 通过多阶段专家模型与丰富提示,自动生成细粒度图像标注
- 在COCO和Visual Genome上使物体标注量翻三倍,描述长度增加15倍
- 适合研究视觉语言模型训练数据增强的学者与工程师
多模态大语言模型(MLLM)在多种视觉-语言任务中展现出强大的推理与泛化能力,但其性能高度依赖监督微调(SFT)阶段的高质量数据。现有方法虽利用GPT-4V构建高质量数据,但受限于GPT-4V的商业属性及提示设计简单,难以规模化。为此,我们提出FullAnno系统——一种可生成大规模、高精度、细粒度图像标注的数据引擎,涵盖物体类别与位置、区域描述、文本信息及图像密集描述。该系统采用多级注释流程,结合多个专家模型与丰富提示,驱动大模型生成密集图像描述。我们使用FullAnno重新标注了COCO和Visual Genome数据集,使物体标注数量提升至原数据的三倍,图像描述长度增长15倍。实验表明,重构后的标注显著提升了LLaVA-v1.5在多个基准测试中的表现。重建数据已公开:https://arcana-project-page.github.io
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have shown promise in a broad range of vision-language tasks with their strong reasoning and generalization capabilities. However, they heavily depend on high-quality data in the Supervised Fine-Tuning (SFT) phase. The existing approaches aim to curate high-quality data via GPT-4V, but they are not scalable due to the commercial nature of GPT-4V and the simplicity of the prompts used to instruct the model. To this end, we devised the FullAnno system, which is a data engine that can generate large-scale, high-quality, and fine-grained image annotations consisting of the category and position of objects, region descriptions, text information, as well as image dense captions. This engine is characterized by its cascade annotation process, which involves multiple expert models and employs rich prompts to instruct LLMs in generating dense image captions. We re-annotated the COCO and Visual Genome datasets using our FullAnno system, tripling the number of object annotations and increasing the length of the original image captions by a factor of 15. Experiments show that the regenerated annotation can significantly enhance the capabilities of LLaVA-v1.5 on several benchmarks. The re-annotated data are available at: https://arcana-project-page.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。