开源低成本训练框架,让普通团队也能造出顶尖多模态模型
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
- 用8500万条数据+2200万条指令,从零开始构建高质量多模态模型
- 在1.6万美元预算内完成训练,4B模型全面超越3B竞品
- 引入轻量强化学习提升复杂推理能力,适合想自主训练的团队
我们提出 LLaVA-OneVision-1.5,一个全新的大型多模态模型家族,在显著降低计算与财务成本的前提下达到业界顶尖性能。不同于现有工作,该框架提供开放、高效且可复现的方案,可完全从零构建高质量视觉语言模型。其发布包含三大核心组件:(1) 大规模精选数据集:构建了8500万条概念均衡的预训练数据集 LLaVA-OneVision-1.5-Mid-Traning 以及2200万条精心标注的指令数据集 LLaVA-OneVision-1.5-Instruct;(2) 高效训练框架:开发端到端高效训练流程,采用离线并行数据打包策略,支持在1.6万美元预算内完成训练;(3) 顶尖性能表现:实验显示,LLaVA-OneVision-1.5 在多个下游任务中表现优异,其中8B版本在27项基准中的18项超越 Qwen2.5-VL-7B,4B版本在全部27项任务上优于 Qwen2.5-VL-3B;(4) 基于强化学习的后训练:通过轻量级强化学习阶段激发模型潜能,有效增强链式思维推理能力,显著提升复杂多模态推理任务表现。
原文摘要 · Abstract (English)
We present LLaVA-OneVision-1.5, a novel family of Large Multimodal Models (LMMs) that achieve state-of-the-art performance with significantly reduced computational and financial costs. Different from the existing works, LLaVA-OneVision-1.5 provides an open, efficient, and reproducible framework for building high-quality vision-language models entirely from scratch. The LLaVA-OneVision-1.5 release comprises three primary components: (1) Large-Scale Curated Datasets: We construct an 85M concept-balanced pretraining dataset LLaVA-OneVision-1.5-Mid-Traning and a meticulously curated 22M instruction dataset LLaVA-OneVision-1.5-Instruct. (2) Efficient Training Framework: We develop a complete end-to-end efficient training framework leveraging an offline parallel data packing strategy to facilitate the training of LLaVA-OneVision-1.5 within a $16,000 budget. (3) State-of-the-art Performance: Experimental results demonstrate that LLaVA-OneVision-1.5 yields exceptionally competitive performance across a broad range of downstream tasks. Specifically, LLaVA-OneVision-1.5-8B outperforms Qwen2.5-VL-7B on 18 of 27 benchmarks, and LLaVA-OneVision-1.5-4B surpasses Qwen2.5-VL-3B on all 27 benchmarks. (4) RL-based Post-training: We unlock the model's latent potential through a lightweight RL stage, effectively eliciting robust chain-of-thought reasoning to significantly boost performance on complex multimodal reasoning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。