arXiv:2510.00040cs.CVcs.AI2025-10被引 1

通过分析模型内在能力,用5%数据实现全量数据效果

Uncovering Intrinsic Capabilities: A Paradigm for Data Curation in Vision-Language Models

  • 基于梯度轨迹无监督发现模型内在能力
  • 仅用5%数据超越全量数据训练效果
  • 适合需要高效微调的AI研发人员

大型视觉语言模型(VLMs)在基准测试中表现强劲,但通过指令微调控制其行为仍具挑战。减少指令微调数据集预算常导致性能下降,因传统启发式策略将模型视为黑箱,忽视了决定学习过程的潜在能力。我们提出能力归因数据整理(CADC)框架,将数据整理从任务特定启发式转向内在能力分析。CADC通过梯度学习轨迹无监督发现内在能力,利用影响估计将训练数据归因于这些能力,并通过均衡选择和分阶段排序构建能力感知的课程。该方法将黑箱指令微调转化为可控制、以能力为导向的过程。仅使用原数据的5%,CADC在多模态基准上即超越全量数据训练效果。结果验证了内在能力是模型学习的基本单元,并确立了CADC作为指令数据整理的范式。

原文摘要 · Abstract (English)

Large vision-language models (VLMs) achieve strong benchmark performance, but controlling their behavior through instruction tuning remains difficult. Reducing the budget of instruction tuning dataset often causes regressions, as heuristic strategies treat models as black boxes and overlook the latent capabilities that govern learning. We introduce Capability-Attributed Data Curation (CADC), a framework that shifts curation from task-specific heuristics to intrinsic capability analysis. CADC discovers intrinsic capabilities in an unsupervised manner from gradient-based learning trajectories, attributes training data to these capabilities via influence estimation, and curates capability-aware curricula through balanced selection and staged sequencing. This transforms black-box instruction tuning into a controllable, capability-driven process. With as little as 5% of the original data, CADC surpasses full-data training on multimodal benchmarks. These results validate intrinsic capabilities as the fundamental building blocks of model learning and establish CADC as a principle paradigm for instruction data curation.

视觉语言模型数据整理指令微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。