在数据受限下研究原生多模态大模型的扩展规律,提出高效新模型NaViL。
NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
- 端到端原生训练,系统探索架构与扩展性设计空间。
- 发现视觉编码器与语言模型规模正相关,提升性能更稳定。
- 适合追求高效训练与实用性能的多模态研究者参考。
现有多模态大模型(MLLM)普遍采用组合式训练范式,即通过连续多模态预训练连接预训练视觉编码器与语言模型。然而,由于训练过程分离,该范式下的多模态扩展特性难以深入探究。本文聚焦于在实际数据约束条件下,以端到端方式开展原生多模态大模型的训练,系统研究其设计空间与扩展特性。通过细致分析多种架构选择,我们获得在性能与训练成本间平衡最优的元架构。进一步研究发现,视觉编码器与语言模型规模存在正相关扩展关系。基于此,提出名为NaViL的原生多模态大模型,并配套简单且低成本的训练方案。14个多模态基准测试结果表明,NaViL性能优于现有主流模型。研究成果为未来原生多模态大模型研究提供了深层洞见。
原文摘要 · Abstract (English)
Compositional training has been the de-facto paradigm in existing Multimodal Large Language Models (MLLMs), where pre-trained vision encoders are connected with pre-trained LLMs through continuous multimodal pre-training. However, the multimodal scaling property of this paradigm remains difficult to explore due to the separated training. In this paper, we focus on the native training of MLLMs in an end-to-end manner and systematically study its design space and scaling property under a practical setting, i.e., data constraint. Through careful study of various choices in MLLM, we obtain the optimal meta-architecture that best balances performance and training cost. After that, we further explore the scaling properties of the native MLLM and indicate the positively correlated scaling relationship between visual encoders and LLMs. Based on these findings, we propose a native MLLM called NaViL, combined with a simple and cost-effective recipe. Experimental results on 14 multimodal benchmarks confirm the competitive performance of NaViL against existing MLLMs. Besides that, our findings and results provide in-depth insights for the future study of native MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。