arXiv:2605.11405cs.LG2026-05被引 1

仅靠数据筛选,20亿参数视觉语言模型性能大幅跃升

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone

论文配图:20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone
图 1 · 摘自论文原文
  • 用精细化数据筛选替代复杂训练,提升模型泛化能力
  • 在20个基准上平均提升11.7个百分点,9项能力轴均显著领先
  • 适合追求高性价比、低算力部署的VLM研发与应用者

数据筛选已显著提升语言模型与对比图像-文本预训练的质量-算力边界,但其对视觉语言模型(VLM)的作用尚不明确。本文在固定架构、训练方法和算力的前提下,仅通过改变训练数据,探究数据筛选对VLM性能的极限影响。基于MAmmoTH-VL单图子集构建的流水线,在20个公开的VLM基准(涵盖定位、VQA、OCR/文档、描述生成、空间/3D、计数、图表、数学、品牌识别及多图推理)上平均提升11.7个百分点;在包含九项能力维度的DatBench评估套件中,平均提升11.3个百分点。20亿参数模型在仅17倍于基线的训练算力下,超越InternVL3.5-2B达9.9个百分点,且与Qwen3-VL-2B差距缩小至1.8个百分点,仅需其87倍算力。此外,数据筛选还带来四大优势:可靠性提升(各能力标准差下降约67%)、跨域泛化增强(9项跨域平均提升7.2个百分点)、行为表现更优(1100个开放问题中更诚实、具体、简洁、拒绝率更低),以及推理成本降低——在1B、2B、4B各规模下,准确率更高同时响应FLOPs更低,4B版本以3.3倍更低的响应算力逼近前沿精度。

原文摘要 · Abstract (English)

Data curation has shifted the quality-compute frontier for language-model and contrastive image-text pretraining, but its role for vision-language models (VLMs) is far less established. We ask how far data curation alone can take VLM performance, holding architecture, training recipe, and compute fixed and varying only the training data. Our pipeline, applied to the MAmmoTH-VL single-image subset, lifts performance by +11.7pp on average across 20 public VLM benchmarks (spanning grounding, VQA, OCR/documents, captioning, spatial/3D, counting, charts, math, brand-ID, and multi-image reasoning) and by +11.3pp on average across all nine capability axes of DatBench, our high-fidelity VLM eval suite. At 2B, our curated model surpasses InternVL3.5-2B by 9.9pp at ~17x less training compute and closes the gap to Qwen3-VL-2B to within 1.8pp at ~87x less compute, from pretraining alone. Beyond accuracy, curation delivers four further properties: (1) Reliability: per-capability std across training seeds drops by ~67% and the lift survives a 4k-to-16k context-length sweep; (2) OOD generalization: the 9-eval OOD average rises by +7.2pp, and multi-image BLINK rises by +3.09pp despite single-image-only training, with Visual Correspondence gaining +11.8pp; (3) Behavioral gains beyond benchmarks: across ~1,100 open-ended queries the curated 2B is more honest and more specific than the matched-compute baseline, and more concise and less refusal-prone than a frontier 2B reference; (4) Pareto-dominance on inference cost: at every scale (1B, 2B, 4B) the curated model raises accuracy while lowering response FLOPs vs. the matched-compute baseline, and the curated 4B matches near-frontier accuracy at 3.3x lower response FLOPs than Qwen3-VL-4B. Data curation is a high-leverage tool for building better VLMs, reaching near-frontier accuracy at up to ~150x less training compute.

视觉语言模型数据筛选算力优化模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。