arXiv:2501.14818cs.CVcs.AI2025-01被引 75

从零构建视觉语言模型训练数据策略,显著提升开源模型性能。

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models

  • 从零设计后训练数据策略,聚焦数据质量与多样性
  • Eagle2-9B在多模态评测中达到顶尖水平,媲美70B参数模型
  • 开源完整数据策略与训练方案,助力社区模型发展

近期,开源视觉语言模型(VLMs)在能力上已接近闭源前沿模型。然而,多数开源模型仅发布最终模型权重,数据策略与实现细节仍不透明。本文从数据驱动视角出发,系统研究并构建了后训练数据策略,揭示其在打造前沿VLM中的关键作用。我们提出一套完整的数据策略,结合训练配方与模型设计,构建出名为Eagle2的高性能模型家族。其中,Eagle2-9B在多个多模态基准测试中表现优异,性能媲美参数高达70B的先进模型,为开源社区提供了可复现、可改进的开发范式。

原文摘要 · Abstract (English)

Recently, promising progress has been made by open-source vision-language models (VLMs) in bringing their capabilities closer to those of proprietary frontier models. However, most open-source models only publish their final model weights, leaving the critical details of data strategies and implementation largely opaque. In this work, we address VLM post-training from a data-centric perspective, showing the key role of data strategy in developing frontier VLMs. By studying and building our post-training data strategy from scratch, we share detailed insights into the development processes, aiming to benefit the development of competitive models for the open-source community. Our introduced data strategy, together with training recipes and model design, leads to a family of performant VLMs named Eagle2. Specifically, Eagle2-9B achieves state-of-the-art results across various multimodal benchmarks, matching certain competitive models with up to 70B parameters.

视觉语言模型数据策略开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。