arXiv:2502.09925cs.CVcs.AI2025-02ICLR被引 8

用自动化方法构建超大规模多模态指令数据集,提升模型泛化能力。

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types

  • 基于GPT-4o与CLIP自动扩展19,227种任务类型
  • 在16个基准上显著提升LLaVA和InternVL性能
  • 适合需要强泛化能力的多模态应用开发者

多模态视觉语言模型在开放世界应用中日益重要,但其表现常受限于任务特定数据不足,导致泛化能力差和输出偏差。现有方法因手动标注耗时,通常仅生成数百种任务类型。为此,我们提出TaskGalaxy,一个包含19,227个层级任务类型和413,648个样本的大规模多模态指令微调数据集。该数据集利用GPT-4o从少量人工定义任务出发进行扩展,结合CLIP和GPT-4o筛选匹配开源图像的任务,并生成相关问答对。多模型协作确保样本质量。该自动化流程显著提升了任务多样性和数据质量,大幅减少人工干预。将TaskGalaxy用于LLaVA-v1.5和InternVL-Chat-v1.0模型,在16个基准测试中均取得显著性能提升,证明任务多样性至关重要。数据集已公开:https://github.com/Kwai-YuanQi/TaskGalaxy。

原文摘要 · Abstract (English)

Multimodal visual language models are gaining prominence in open-world applications, driven by advancements in model architectures, training techniques, and high-quality data. However, their performance is often limited by insufficient task-specific data, leading to poor generalization and biased outputs. Existing efforts to increase task diversity in fine-tuning datasets are hindered by the labor-intensive process of manual task labeling, which typically produces only a few hundred task types. To address this, we propose TaskGalaxy, a large-scale multimodal instruction fine-tuning dataset comprising 19,227 hierarchical task types and 413,648 samples. TaskGalaxy utilizes GPT-4o to enrich task diversity by expanding from a small set of manually defined tasks, with CLIP and GPT-4o filtering those that best match open-source images, and generating relevant question-answer pairs. Multiple models are employed to ensure sample quality. This automated process enhances both task diversity and data quality, reducing manual intervention. Incorporating TaskGalaxy into LLaVA-v1.5 and InternVL-Chat-v1.0 models shows substantial performance improvements across 16 benchmarks, demonstrating the critical importance of task diversity. TaskGalaxy is publicly released at https://github.com/Kwai-YuanQi/TaskGalaxy.

多模态指令微调数据集自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。