从零构建视觉语言模型训练数据策略,显著提升开源模型性能。
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
- 从零设计后训练数据策略,聚焦数据质量与多样性
- Eagle2-9B在多模态评测中达到顶尖水平,媲美70B参数模型
- 开源完整数据策略与训练方案,助力社区模型发展
近期,开源视觉语言模型(VLMs)在能力上已接近闭源前沿模型。然而,多数开源模型仅发布最终模型权重,数据策略与实现细节仍不透明。本文从数据驱动视角出发,系统研究并构建了后训练数据策略,揭示其在打造前沿VLM中的关键作用。我们提出一套完整的数据策略,结合训练配方与模型设计,构建出名为Eagle2的高性能模型家族。其中,Eagle2-9B在多个多模态基准测试中表现优异,性能媲美参数高达70B的先进模型,为开源社区提供了可复现、可改进的开发范式。
原文摘要 · Abstract (English)
Recently, promising progress has been made by open-source vision-language models (VLMs) in bringing their capabilities closer to those of proprietary frontier models. However, most open-source models only publish their final model weights, leaving the critical details of data strategies and implementation largely opaque. In this work, we address VLM post-training from a data-centric perspective, showing the key role of data strategy in developing frontier VLMs. By studying and building our post-training data strategy from scratch, we share detailed insights into the development processes, aiming to benefit the development of competitive models for the open-source community. Our introduced data strategy, together with training recipes and model design, leads to a family of performant VLMs named Eagle2. Specifically, Eagle2-9B achieves state-of-the-art results across various multimodal benchmarks, matching certain competitive models with up to 70B parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。