arXiv:2608.05000cs.CVcs.LG2026-08被引 1

揭示多模态预训练中知识流动、模态协同与早期融合的内在机制。

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

论文配图:Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
图 1 · 摘自论文原文
  • 通过控制实验分离语言与视觉间知识传递路径,发现不对称影响模式。
  • 早期联合训练比后期对齐更有效,延迟融合会引发视觉依赖语言先验。
  • 提出高效训练方案,仅用5%算力即达强生成性能,适用于大规模模型。

视觉为推动基础模型发展提供了关键维度,推动了原生统一的多模态预训练趋势。尽管如此,多模态统一训练中的设计空间与模态交互机制仍缺乏深入探索。本文通过在合成数据和大规模真实数据集上的系统性实验,揭示多模态预训练的四大核心洞见:(i) 知识流动:解耦语言、视觉理解与视觉生成之间的知识迁移,揭示其影响模式与不对称性;(ii) 协同与竞争:数据“复杂度”决定模态是否协同,共享注意力与模态特定前馈层可促进协同,且该行为在不同视觉分词器设计下具泛化性;(iii) 早期统一:从训练初期即联合统一模态优于后期对齐或顺序训练,揭示“视觉惰性”现象——延迟整合导致模型过度依赖语言先验;(iv) 训练配方:提出高效预训练方案,仅使用5%算力即可实现强大生成性能。上述发现经由在2T token上训练多个13.5B MoE模型的规模化验证,为理解与扩展多模态预训练提供原则性基础。

原文摘要 · Abstract (English)

Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining. Despite this momentum, the design space and the fundamental mechanisms of how modalities interact during unified training remain underexplored. We provide empirical clarity through a systematic exploration of multimodal pretraining. Our controlled experiments on both synthetic and large-scale real-world datasets yield four key insights into the physics of multimodal pretraining: (i) Knowledge Flow: We disentangle how language, visual understanding, and visual generation transfer knowledge across modalities, revealing distinct patterns of influence and asymmetry; (ii) Synergy vs. Competition: We show that data "complexity" largely determines whether modalities are synergistic, identify architectural choices that promote synergy: such as shared attention and normalization with modality-specific feed-forward layers, and find that these behaviors generalize across different visual tokenizer designs; (iii) Early Unification: Unifying modalities from the very early stages and training them jointly is shown to be more effective than late alignment or sequential training. This process uncovers a vision laziness phenomenon, where delayed integration leads models to rely on language priors; (iv) Recipes: We derive efficient pretraining recipes that achieve strong generative performance using only 5% of the compute budget. These core findings are subsequently validated at scale by training multiple 13.5B MoE models on 2T tokens. We hope this study provides a principled foundation for understanding and scaling multimodal pretraining.

多模态预训练知识流动模型架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。