arXiv:2410.04038cs.AIcs.CV2024-10被引 1

用游戏化方式收集高质量视觉指令数据,显著提升小模型表现

Gamified crowd-sourcing of high-quality data for visual fine-tuning

  • 将数据收集设计为游戏,激励用户提出针对模型弱点的细粒度问答
  • 在数周内从5万+参与者获取数据,使小模型GPT评分从0.147升至0.477
  • 生成的数据不仅提升本模型,还通用增强其他多模态模型性能

本文提出游戏化对抗提示(GAP)框架,通过游戏化机制众包高质量视觉指令数据,用于大模型微调。GAP将数据收集转化为互动游戏,激励玩家提供能暴露模型知识盲区的精细问题与答案。该方法包含三项贡献:(1)直接捕捉针对模型缺陷的问答对;(2)评估并奖励高质量提交的机制;(3)可扩展的游戏化平台,仅用数周即从5万余名参与者中收集数据。实验表明,使用GAP生成的数据使小型多模态模型MiniCPM-Llama3-V-2.5-8B在自建数据集上的GPT评分从0.147提升至0.477,接近大型模型GPT-4V的表现。此外,该数据还跨模型提升QWEN2-VL-2B和QWEN2-VL-7B在多个基准上的性能。

原文摘要 · Abstract (English)

This paper introduces Gamified Adversarial Prompting (GAP), a framework that crowd-sources high-quality data for visual instruction tuning of large multimodal models. GAP transforms the data collection process into an engaging game, incentivizing players to provide fine-grained, challenging questions and answers that target gaps in the model's knowledge. Our contributions include (1) an approach to capture question-answer pairs from humans that directly address weaknesses in a model's knowledge, (2) a method for evaluating and rewarding players that successfully incentivizes them to provide high-quality submissions, and (3) a scalable, gamified platform that succeeds in collecting this data from over 50,000 participants in just a few weeks. Our implementation of GAP has significantly improved the accuracy of a small multimodal model, namely MiniCPM-Llama3-V-2.5-8B, increasing its GPT score from 0.147 to 0.477 on our dataset, approaching the benchmark set by the much larger GPT-4V. Moreover, we demonstrate that the data generated using MiniCPM-Llama3-V-2.5-8B also enhances its performance across other benchmarks, and exhibits cross-model benefits. Specifically, the same data improves the performance of QWEN2-VL-2B and QWEN2-VL-7B on the same multiple benchmarks.

多模态数据众包游戏化模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。