通过数据代谢框架,用高效数据设计让小模型超越大模型。
Data Metabolism: An Efficient Data Design Schema For Vision Language Model
- 构建闭环数据迭代系统,持续优化视觉语言模型训练数据。
- 小型模型Capybara-VL性能超越10倍大的开源模型,接近顶级闭源模型。
- 适合追求高效、低成本训练的模型研发团队使用。
数据筛选在训练强大视觉语言模型(VLMs)中起着关键作用。本文提出‘数据代谢’概念,构建以数据为中心的框架,贯穿模型开发全生命周期。从标准模型架构出发,深入探讨数据筛选与迭代两个关键步骤,形成持续提升模型性能的闭环系统。我们提供详细的数据处理手册,用于整合现有大规模数据集并构建用户专属的数据飞轮。作为演示,我们发布一个名为Capybara-VL的VLM,在典型多模态任务(如视觉问答、科学推理和文本密集型任务)中表现优异。尽管模型规模相对较小,其性能仍超越多个规模高达10倍的开源模型,并达到若干领先闭源模型的水平,凸显了该数据驱动框架的强大潜力及训练更小、更高效VLM的可能性。
原文摘要 · Abstract (English)
Data curation plays a crucial role in training powerful Visual Language Models (VLMs). In this work, we introduce the concept of Data Metabolism and present our data-centric framework to build VLMs throughout the development lifecycle. Starting from a standard model architecture, we discuss and provide insights into two crucial development steps: data curation and iteration, forming a closed-loop system that continuously improves model performance. We show a detailed codebook on how to process existing massive datasets and build user-specific data flywheel. As a demonstration, we release a VLM, named Capybara-VL, which excels in typical multimodal tasks (e.g. , visual question answering, scientific reasoning, and text-rich tasks). Despite its relatively compact size, Capybara-VL surpasses several open-source models that are up to 10 times larger in size. Moreover, it achieves results that are on par with those of several leading proprietary models, demonstrating its remarkable competitiveness. These results highlight the power of our data-centric framework and the potential of training smaller and more efficient VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。