arXiv:2607.22043cs.CLcs.CV2026-07被引 1

提出可预测的多模态模型扩展规律,指导高效训练。

Scaling Native Multimodal Pre-Training From Scratch

  • 在固定算力下优化模型规模与数据量,实现高效多模态预训练。
  • 语言任务学习稳定,多模态任务效率随文本比例提升而变化。
  • 给出资源分配最优解,适用于追求高效多模态模型的研究者。

尽管大语言模型具备出色推理能力,但其仅依赖文本预训练限制了对多模态物理世界的感知。原生多模态预训练通过从零开始在多模态输入上训练模型,实现深度跨模态融合,并缓解传统晚期融合架构中的优化不对称问题。然而,该范式下的扩展特性尚未系统研究。本文在固定计算预算下,探究基于Transformer的视觉-语言模型的最佳模型规模与标记数量。结果表明,最小目标损失遵循可预测的算力定律,而算力最优的模型规模与标记数量呈幂律增长。值得注意的是,语言与多模态目标展现出不同的扩展行为:语言分配法则几乎不受数据构成影响,表明语言学习稳定;而多模态分配法则高度依赖数据组成。具体而言,以文本为主的混合数据仅在更大模型规模下才具计算效率,导致最优资源配置向更高模型容量倾斜。此外,通过建模数据构成对算力定律和分配指数的影响,推导出精确的效率前沿,明确指定模型规模、标记数量与数据混合比的最优配置。下游评估进一步显示,原生多模态预训练带来正向跨模态迁移,提升纯文本空间推理能力并实现鲁棒的多模态上下文学习。综上,本实证研究为可预测地扩展多模态基础模型奠定了关键基础。

原文摘要 · Abstract (English)

Although large language models (LLMs) exhibit remarkable reasoning capabilities, their reliance on text-only pre-training restricts the perception of the multimodal physical world. Native multimodal pre-training avoids this limitation by training models from scratch on multimodal inputs, thereby achieving deep cross-modal integration and mitigating optimization asymmetries inherent to traditional late-fusion architectures. Despite these advantages, the scaling properties of this paradigm remain systematically uncharacterized. To address this gap, we investigate the optimal model size and token count for training a transformer-based vision-language model under a fixed computational budget. We demonstrate that minimal objective loss adheres to a predictable compute law, whereas compute-optimal model sizes and token counts scale as power laws. Notably, language and multimodal objectives manifest distinct scaling behaviors. The language allocation law is largely invariant to the composition of the data, indicating stable language learning regardless of the multimodal data ratio. Conversely, the multimodal allocation law is highly sensitive to this composition. Specifically, text-heavy mixtures become compute-efficient only at larger model scales, shifting the optimal resource allocation toward greater model capacity. Additionally, by modeling the influence of data composition on compute laws and allocation exponents, we derive an efficiency frontier specifying precise configurations of model size, token count, and data mixture. Downstream evaluations further reveal that native multimodal pre-training induces positive cross-modal transfer, thereby enhancing pure-text spatial reasoning and enabling robust multimodal in-context learning. In summary, this empirical research establishes the essential groundwork for predictably scaling multimodal foundation models.

多模态预训练扩展规律模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。