arXiv:2503.13383cs.CVcs.AI2025-03被引 1

提出高效筛选多模态指令数据的方法,仅用30%数据达到99.1%性能。

Cream of the Crop: Harvesting Rich, Scalable and Transferable Multi-Modal Data for Instruction Fine-Tuning

  • 按14项视觉语言能力拆解质量评估,精准识别优质数据
  • 通过交互风格实现多样性控制,避免数据模式单一
  • 无需嵌入聚类或贪心采样,支持百万级数据快速处理

预训练大模型在微调阶段仅需少量监督的假设已得到验证,但其稳定性与泛化能力受限于实验设置和验证协议,表现不如随机采样。多模态大模型因数据量大、来源异质性强,数据选择更关键也更复杂。为此,本文将质量指标细分为14项视觉-语言相关能力,引入多模态丰富评分器(mmSSR)评估每条数据能力;为提升多样性,以交互风格为指标,使用多模态丰富风格器识别指令模式。mmSSR确保高分信息以多样化形式呈现。该方法无需基于嵌入的聚类或贪心采样,可高效扩展至百万级数据,支持不同预算和能力目标定制,并实现无需训练即可迁移至新领域。在超过10组实验设置下,经14个多模态基准验证,相较随机采样、基线及现有最优方法均有持续提升,在仅使用260万数据的30%(即78万)时,达到全量数据99.1%的性能。

原文摘要 · Abstract (English)

The hypothesis that pretrained large language models (LLMs) necessitate only minimal supervision during the fine-tuning (SFT) stage (Zhou et al., 2024) has been substantiated by recent advancements in data curation and selection research. However, their stability and generalizability are compromised due to the vulnerability to experimental setups and validation protocols, falling short of surpassing random sampling (Diddee & Ippolito, 2024; Xia et al., 2024b). Built upon LLMs, multi-modal LLMs (MLLMs), combined with the sheer token volume and heightened heterogeneity of data sources, amplify both the significance and complexity of data selection. To harvest multi-modal instructional data in a robust and efficient manner, we re-define the granularity of the quality metric by decomposing it into 14 vision-language-related capabilities, and introduce multi-modal rich scorers to evaluate the capabilities of each data candidate. To promote diversity, in light of the inherent objective of the alignment stage, we take interaction style as diversity indicator and use a multi-modal rich styler to identify data instruction patterns. In doing so, our multi-modal rich scorers and styler (mmSSR) guarantee that high-scoring information is conveyed to users in diversified forms. Free from embedding-based clustering or greedy sampling, mmSSR efficiently scales to millions of data with varying budget constraints, supports customization for general or specific capability acquisition, and facilitates training-free generalization to new domains for curation. Across 10+ experimental settings, validated by 14 multi-modal benchmarks, we demonstrate consistent improvements over random sampling, baseline strategies and state-of-the-art selection methods, achieving 99.1% of full performance with only 30% of the 2.6M data.

多模态数据筛选指令微调高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。