针对多模态模型指令选择,区分概念与技能并针对性选数据,提升性能。
Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models
- 按任务需求筛选匹配概念或技能的指令数据
- 平均提升0.9%,技能类任务提升1.5%
- 适合希望优化多模态模型训练效率的研究者
视觉-语言指令微调主要服务于学习视觉概念和视觉技能。本文发现,现有视觉-语言基准测试的表现主要依赖于相似技能或视觉概念的训练指令。基于此,我们提出一种简单的目标导向数据选择方法:先提取基准中的概念/技能,判断其更依赖概念或技能,再选取最匹配的概念/技能的指令。在10+个基准上的实验验证了该方法的有效性,平均性能优于最佳基线0.9%,在技能类任务上提升1.5%。研究强调了指令选择中概念知识与视觉技能之间的内在权衡,需根据任务特性进行平衡。
原文摘要 · Abstract (English)
Vision-language instruction tuning achieves two main purposes: learning visual concepts and learning visual skills. In this paper, we found that vision-language benchmarks fall into the dichotomy of mainly benefiting from training on instructions with similar skills or visual concepts. Inspired by the discovery, we designed a simple targeted training data selection method to optimize the performance of a given benchmark. We first extract the concepts/skills from the benchmark, determine whether the benchmark predominantly benefits from similar concepts or skills, and finally select instructions with the most matching concepts/skills. Experiments on 10+ benchmarks validate the effectiveness of our targeted data selection method, showing +0.9\% over the best existing baseline averaged over all benchmarks and +1.5\% on the skill-focused subset. Our findings underscore the importance of recognizing the inherent trade-off within instruction selection, which requires balancing the acquisition of conceptual knowledge against visual skill.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。