arXiv:2504.21850cs.CV2025-04被引 1

通过组合视觉能力生成复杂训练样本,大幅降低视觉指令微调数据量。

Visual Compositional Tuning

  • 将多个基础视觉能力合成到单一样本中,提升单条数据信息密度。
  • 仅用10%原始数据量,实现97.5%以上性能(原为100.2%)。
  • 适合追求高效训练的视觉语言模型研究者与开发者。

视觉指令微调(VIT)数据集规模迅速增长,但单个训练样本的信息量常被忽视。近期研究表明,少量富含信息的样本即可实现多模态大模型的有效微调。本文探索样本复杂度对数据筛选的影响,提出COMPACT(COMPositional Atomic-to-complex Visual Compositional Tuning)方法,通过在单个训练样本中组合多种基础视觉能力,提升样本复杂度。具体地,为每张图像生成丰富且信息量高的文本问题,显著减少所需训练样本数量。在LLaVA-665K数据集上,COMPACT将数据预算降低90%,仍保持100.2%的完整微调性能(优于当前最优方法的97.5%),并在八项多模态基准测试中表现优异。尤其在复杂任务如MM-Vet(+8.6%)和MMStar(+2.9%)上,其性能超越全量数据训练。COMPACT提供了一种可扩展、高效的合成数据生成方案,用于提升视觉语言任务性能。

原文摘要 · Abstract (English)

Visual instruction tuning (VIT) datasets have grown rapidly in scale, yet the informativeness of individual training samples has largely been overlooked. Recent dataset selection methods have shown that a small fraction of such datasets enriched with informative samples can lead to efficient finetuning of Multimodal Large Language Models. In this work, we explore the impact of sample complexity on informative data curation and introduce COMPACT (COMPositional Atomic-to-complex Visual Compositional Tuning), a compositional VIT data recipe that scales training sample complexity by combining multiple atomic visual capabilities in a single training example. Concretely, we synthesize rich and informative text questions for each image, allowing us to significantly reduce the number of training examples required for effective VIT. COMPACT demonstrates superior data efficiency compared to existing data reduction methods. When applied to the LLaVA-665K VIT dataset, COMPACT reduces the data budget by 90% while still achieving 100.2% of the full VIT performance (compared to only 97.5% by the state-of-the-art method) across eight multimodal benchmarks. Furthermore, training on the COMPACT data outperforms training on the full-scale VIT data on particularly complex benchmarks such as MM-Vet (+8.6%) and MMStar (+2.9%). COMPACT offers a scalable and efficient synthetic data generation recipe to improve on vision-language tasks.

视觉语言数据效率合成数据指令微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。