arXiv:2502.07617cs.CV2025-02被引 26

用1000亿数据训练视觉语言模型,提升文化多样性与低资源语言表现。

Scaling Pre-training to One Hundred Billion Data for Vision Language Models

  • 在1000亿级网页数据上预训练视觉语言模型。
  • 文化多样性任务显著受益,低资源语言性能提升明显。
  • 去噪过滤可能削弱数据文化多样性,需权衡使用。

我们对在空前规模(1000亿样本)下预训练视觉语言模型的潜力进行了实证研究。发现模型在多个以西方为中心的分类与检索基准(如COCO Captions)上性能趋于饱和。然而,涉及文化多样性的任务从1000亿规模的网络数据中获得了更显著的提升,得益于其对长尾概念的覆盖。此外,我们分析了模型的多语言能力,发现低资源语言性能也有提升。值得注意的是,通过CLIP等质量过滤器减少预训练数据量虽常用于提升性能,但可能无意中降低大规模数据集中的文化多样性。结果表明,尽管传统基准在扩展至1000亿样本后收益有限,但该数据规模对构建真正包容的多模态系统至关重要。

原文摘要 · Abstract (English)

We provide an empirical investigation of the potential of pre-training vision-language models on an unprecedented scale: 100 billion examples. We find that model performance tends to saturate at this scale on many common Western-centric classification and retrieval benchmarks, such as COCO Captions. Nevertheless, tasks of cultural diversity achieve more substantial gains from the 100-billion scale web data, thanks to its coverage of long-tail concepts. Furthermore, we analyze the model's multilinguality and show gains in low-resource languages as well. In addition, we observe that reducing the size of the pretraining dataset via quality filters like using CLIP, typically used to enhance performance, may inadvertently reduce the cultural diversity represented in large-scale datasets. Our results highlight that while traditional benchmarks may not benefit significantly from scaling noisy, raw web data to 100 billion examples, this data scale is vital for building truly inclusive multimodal systems.

视觉语言模型多模态数据规模文化多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。