arXiv:2508.10339cs.CVcs.LG2025-08被引 1

针对多模态模型指令选择,区分概念与技能并针对性选数据,提升性能。

Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models

  • 按任务需求筛选匹配概念或技能的指令数据
  • 平均提升0.9%,技能类任务提升1.5%
  • 适合希望优化多模态模型训练效率的研究者

视觉-语言指令微调主要服务于学习视觉概念和视觉技能。本文发现,现有视觉-语言基准测试的表现主要依赖于相似技能或视觉概念的训练指令。基于此,我们提出一种简单的目标导向数据选择方法:先提取基准中的概念/技能,判断其更依赖概念或技能,再选取最匹配的概念/技能的指令。在10+个基准上的实验验证了该方法的有效性,平均性能优于最佳基线0.9%,在技能类任务上提升1.5%。研究强调了指令选择中概念知识与视觉技能之间的内在权衡,需根据任务特性进行平衡。

原文摘要 · Abstract (English)

Vision-language instruction tuning achieves two main purposes: learning visual concepts and learning visual skills. In this paper, we found that vision-language benchmarks fall into the dichotomy of mainly benefiting from training on instructions with similar skills or visual concepts. Inspired by the discovery, we designed a simple targeted training data selection method to optimize the performance of a given benchmark. We first extract the concepts/skills from the benchmark, determine whether the benchmark predominantly benefits from similar concepts or skills, and finally select instructions with the most matching concepts/skills. Experiments on 10+ benchmarks validate the effectiveness of our targeted data selection method, showing +0.9\% over the best existing baseline averaged over all benchmarks and +1.5\% on the skill-focused subset. Our findings underscore the importance of recognizing the inherent trade-off within instruction selection, which requires balancing the acquisition of conceptual knowledge against visual skill.

多模态指令微调数据选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。