构建42.4万对高质量图文指令数据,提升模型对含文字图像的理解能力。
A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation
- 融合人工标注与GPT-4o生成,通过精细化提示工程增强图文对齐
- 产出42.4万组高质量指令-图像对,显著优于自动生成数据
- 适合训练视觉语言模型处理含文字的复杂图像,如广告、界面等
大型多模态模型在处理含文字图像时仍面临挑战,主要因训练数据不足。自动生成指令虽无需标注,但质量较差,即使最大模型也难以实现良好图文对齐。本文提出LLaVAR-2,通过结合人工标注的详细图像描述与针对GPT-4o定制的文本提示,实现混合式指令生成。该方法引入多项过滤机制以剔除低质数据,最终构建出包含42.4万对高质量指令-图像样本的数据集。实证结果表明,基于该数据集微调的模型,在性能上显著超越使用自生成数据训练的模型。
原文摘要 · Abstract (English)
Large multimodal models still struggle with text-rich images because of inadequate training data. Self-Instruct provides an annotation-free way for generating instruction data, but its quality is poor, as multimodal alignment remains a hurdle even for the largest models. In this work, we propose LLaVAR-2, to enhance multimodal alignment for text-rich images through hybrid instruction generation between human annotators and large language models. Specifically, it involves detailed image captions from human annotators, followed by the use of these annotations in tailored text prompts for GPT-4o to curate a dataset. It also implements several mechanisms to filter out low-quality data, and the resulting dataset comprises 424k high-quality pairs of instructions. Empirical results show that models fine-tuned on this dataset exhibit impressive enhancements over those trained with self-instruct data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。